
The failure that felt normal
A client's AI image feature failed on the same schedule every day since launch. The team called it normal and quietly lost every user who hit it.

A client's AI image feature failed on the same schedule every day since launch. The team called it normal and quietly lost every user who hit it.

A pattern for AI agents: make findings the unit of output, gate top-severity claims on objective evidence, and require cross-source corroboration.

How to predict what PostgreSQL would do with a query without running it: the statistics the planner reads and where an offline analyzer finds them.

Configurable trace sampling, per-process opt-out for high-volume tiers, a version-independent OpenSSL path, and graceful degradation on old kernels.

A redrawn logs page with collapsible filters, Lighthouse on its own tab, a self-revising dashboard agent, and packet drop and DNS shape recording.

More context around a slow query: concurrent requests matched by interval overlap, an offline SQL Explain analyzer, and a cross-pod comparison view.

AI-generated dashboards with a visualization linter and model fallback, three hand-built dashboards, and hybrid full-text log search via Manticore.

A tree-search investigation mode, dedicated ingress-controller dashboards with throttling and scheduling delay, and synthetic URL monitoring.

Lower GC, network, eBPF, and memory cost on the host, plus visibility into Unix sockets, PostgreSQL schemas, Redis over TLS, and same-node traffic.

Notes from a Tel Aviv meetup: why developer skepticism is mostly outdated, where the real concerns sit, and how adoption spreads without a mandate.

Design notes from an incident investigation agent: a tree of hypotheses, specialized subagents, evidence-scored evaluation, and bounded search.

A hands-on Claude Code session: structuring prompts, holding context across a large codebase, and the habits that stop you burning tokens.

Once the layers and the instrumentation are in place, on-call changes. The practices that cut escalations from the rotation to the dev team.

Observability is not a tool you buy, it is code. How much of a production system is instrumentation, and what that costs in developer time.

An observability strategy built from the users a system serves: six layers from user experience down to infrastructure, and what each one answers.