Examples

Incident console

The operations UI you would otherwise write by hand — assembled from tracing, metrics, issues and one query per panel.

The overview panel

One call answers “is anything on fire”. Cheap enough to poll every few seconds because it reads a bounded in-process buffer.

yatta/backend/ops/overview.tsts
import { createAPI } from "yatta/api";import { observer } from "../../func/observe"; const route = createAPI("/ops"); route.get("/overview", async () => {  // getDeepApplicationState is async because it reads live subsystem metrics.  const o = await observer.getDeepApplicationState();   // Derive a verdict from the numbers, and say what it is based on.  const errorRate = o.apdex.total > 0 ? o.errors.count / o.apdex.total : 0;   return Response.json({    service: o.service,    environment: o.environment,    uptimeSeconds: o.uptimeSeconds,     health: {      status:        o.errors.count > 0 || o.apdex.p95 > 1000 ? "degraded" : "healthy",      basedOn: `${o.apdex.total} recent request(s)`,    },     latency: { p50: o.apdex.p50, p95: o.apdex.p95, p99: o.apdex.p99, apdex: o.apdex.apdex },    errorRate,    issues: o.errors,    subsystems: { database: o.database, queues: o.queues, cache: o.cache },  });});

The issue queue

yatta/backend/ops/issues.tsts
route.get("/issues", async (ctx) => {  const status = ctx.query().status ?? "unresolved";   // issues is the live map; filter rather than reaching for a convenience  // accessor that does not exist.  const all = [...observer.errors.issues.values()];  const issues =    status === "all" ? all : all.filter((i) => i.status === "unresolved");   return Response.json({    stats: observer.errors.stats(),    issues: issues.map((i) => ({      fingerprint: i.fingerprint,      status: i.status,      name: i.name,      message: i.message,      count: i.count,      usersAffected: i.usersAffected,      eventsPerHour: i.eventsPerHour,      firstSeen: i.firstSeen,      lastSeen: i.lastSeen,      // "Which deploy broke this" is the question that ends an incident.      releases: i.releases,      environments: i.environments,      routes: i.routes,    })),  });}); /** Triage actions, matching the issue lifecycle. */for (const action of ["resolve", "ignore", "unresolve"] as const) {  route.post(`/issues/:fingerprint/${action}`, async (ctx) => {    const by = ctx.query().by ?? "unknown";    const issue = observer.errors.applyAction(ctx.params.fingerprint, {      type: action,      by,    });     if (!issue) throw new HttpError(404, "Unknown issue");     // Triage is itself an audit event.    await record({      tenantId: null,      actorId: by,      action: `issue.${action}`,      target: `issue:${issue.fingerprint}`,      before: null,      after: { status: issue.status },    });     return Response.json({ status: issue.status });  });}

Issue to trace to span

An issue links to the trace it came from, a trace links to its spans, and a span links to the logs written inside it. That chain is the whole debugging workflow.

yatta/backend/ops/trace.tsts
route.get("/traces/:id", async (ctx) => {  const spans = observer.spans.all().filter((sp) => sp.traceId === ctx.params.id);  if (spans.length === 0) throw new HttpError(404, "Unknown trace");   // Assemble the tree the server sends flat.  const byId = new Map(spans.map((s) => [s.spanId, { ...s, children: [] as string[] }]));  let root = spans[0].spanId;   for (const span of spans) {    if (!span.parentSpanId) { root = span.spanId; continue; }    byId.get(span.parentSpanId)?.children.push(span.spanId);  }   // Waterfall offsets, relative to the trace start.  const start = Math.min(...spans.map((s) => s.startTime));   return Response.json({    traceId: ctx.params.id,    rootSpanId: root,    startedAt: start,    totalMs: Math.max(...spans.map((s) => s.endTime ?? s.startTime)) - start,    spans: spans.map((s) => ({      ...s,      offsetMs: s.startTime - start,      durationMs: (s.endTime ?? Date.now()) - s.startTime,      depth: depthOf(s.spanId, byId),      // The link back to an issue, when this span produced one.      issue: observer.errors        .filter((i) => i.occurrences.some((o) => o.spanId === s.spanId))        .map((i) => i.fingerprint),    })),  });}); function depthOf(spanId: string, byId: Map<string, { parentSpanId?: string }>): number {  let depth = 0;  let cursor = byId.get(spanId)?.parentSpanId;  while (cursor && depth < 32) { depth++; cursor = byId.get(cursor)?.parentSpanId; }  return depth;}

Deploy correlation

The single most useful panel during an incident: everything that happened, on one time axis.

yatta/backend/ops/timeline.tsts
route.get("/timeline", async (ctx) => {  const hours = Math.min(Number(ctx.query().hours ?? 6), 168);  const since = Date.now() - hours * 3_600_000;   const issues = [...observer.errors.issues.values()]    .filter((i) => i.lastSeen > since)    .sort((a, b) => b.lastSeen - a.lastSeen)    .slice(0, 200);   const traces = observer.spans.all().filter((sp) => sp.startTime > since);   const events = [    // Deploys come from the release field, not a separate tracker.    ...releasesSince(since).map((r) => ({      at: r.deployedAt,      kind: "deploy" as const,      label: `Deploy ${r.release}`,      detail: r.author,    })),     ...issues.map((i) => ({      at: i.firstSeen,      kind: "issue" as const,      label: `${i.name}: ${i.message.slice(0, 80)}`,      detail: `${i.count} events`,      fingerprint: i.fingerprint,    })),     // Slow requests, bucketed so the panel is readable.    ...bucketLatency(traces, since).map((b) => ({      at: b.at,      kind: "latency" as const,      label: `p95 ${b.p95}ms`,      detail: `${b.count} traces`,    })),     // AI spend, if any.    ...observer.spans.all()      .filter((s) => (s.startTime ?? 0) > since && (s.ai?.costUsd ?? 0) > 0)      .map((s) => ({        at: s.startTime ?? 0,        kind: "ai" as const,        label: `${s.ai?.model}: ${s.ai?.inputTokens}+${s.ai?.outputTokens} tokens`,        detail: `$${(s.ai?.costUsd ?? 0).toFixed(4)}`,        traceId: s.traceId,      })),  ].sort((a, b) => a.at - b.at);   return Response.json({ since, events });});
Tip
Putting deploys on the same axis as errors ends most incidents faster than any dashboard: if the first error follows the newest release by four minutes, you already know where to look.

Alerting

yatta/func/alerts.tsts
export interface AlertRule {  name: string;  /** Evaluated every 60s against the in-process snapshot. */  test: () => Promise<boolean>;  severity: "page" | "ticket";  notify: string[];} export const ALERTS: AlertRule[] = [  {    name: "unhandled-errors",    severity: "page",    test: async () => (await observer.getDeepApplicationState()).errors.count > 0,    notify: ["pagerduty"],  },  {    name: "latency-p95",    severity: "page",    test: async () => (await observer.getDeepApplicationState()).apdex.p95 > 2000,    notify: ["pagerduty", "slack"],  },  {    name: "exporter-failing",    severity: "ticket",    // You are blind if telemetry is not leaving, so this matters even though    // it says nothing about the app.    test: async () => (await observer.getDeepApplicationState()).slowQueries.length > 3,    notify: ["slack"],  },  {    name: "ai-spend",    severity: "ticket",    test: async () => aiSpendToday() > 500,    notify: ["slack"],  },]; /** One cron evaluates every rule; a failing rule must not stop the others. */export const cron = createCron(jobs); cron.schedule("* * * * *", async () => {  for (const rule of ALERTS) {    try {      const firing = await rule.test();      const wasFiring = firingState.has(rule.name);      if (firing === wasFiring) continue;   // only on edges       firingState.set(rule.name, firing);       const snapshotMetrics = async () => {        const o = await observer.getDeepApplicationState();        return { errorCount: o.errors.count, p95: o.apdex.p95, apdex: o.apdex.apdex };      };       await jobs.enqueue("alert:dispatch", {        name: rule.name,        severity: rule.severity,        state: firing ? "firing" : "resolved",        snapshot: await snapshotMetrics(),      });    } catch (err) {      observer.log.error("Alert rule failed", { rule: rule.name, err });    }  }});
Warning
Alert on edges, not on state. A rule that fires every minute while a condition is true is how an alert channel becomes something everyone mutes.