Ketan Khairnar

Systems writing for AI agents, data platforms, distributed coordination, and the production details tutorials skip.

AI-native systems engineering

demo clean production messy

Notes on building systems that hold up in production.

Notes on retries, state, budgets, coordination, observability, and failure recovery — the details that decide whether pressure becomes progress or repeat pain.

distributed systems production agents data platforms observability
10B+ events/day systems direction
150+ paying users from zero
70+ data pipelines shipped
1B+ market data points served

Search

Find the exact thread.

Jump straight into the essays, concepts, explainers, and deep dives behind the system you are thinking about.

agents

AI systems that survive real users

Multi-agent orchestration, bounded tool calls, memory, evals, cost guardrails, and recovery paths.

systems

Distributed systems with operational taste

Coordination, retries, idempotency, stream processing, failure detection, SLOs, and quiet incident prevention.

platforms

Data platforms that earn their keep

Kafka, Spark, ClickHouse, search, lakehouse migrations, and the product pressure behind architecture choices.

Latest field note

Continue reading

Sep 17, 2026
Seven Checks for Systems That Look Correct A May–September retrospective reading path through data correctness, Kafka, RAG, caching, MCP authorization and evaluation, with refreshed series chapters. Seven checks connecting refreshed data, Kafka and AI series. systemsAIdatareading-guide