The runtime

Observability

Tracing, metrics, logs, issue grouping, incident correlation, release intelligence, adaptive baselines, SLOs, N+1 detection and golden traces — in-process.

yatta/observe is built into the framework. One Observer records spans, metrics, logs and issues, and derives analysis from them. Spans propagate through AsyncLocalStorage, so a trace follows a request into a job worker and out over SSE without passing anything around. Nothing leaves the process.

Note
Everything on this page is a current API. Where a capability is not implemented, it is not documented here — check the docs index for the live surface.

Setup

yatta/func/observe.tsts
import { createObserver } from "yatta/observe"; export const observer = createObserver({  service: "app",  release: process.env.GIT_SHA,  // Logical OR, not ?? — an empty NODE_ENV is common in Docker and CI, and a  // nullish coalesce treats "" as set, so a production container would read as  // development and serve the dashboard.  environment: process.env.NODE_ENV || "development",   bufferSize: 2000,   // Everything below this level is dropped before it is masked, buffered,  // breadcrumb'd or broadcast. “silent” turns logging off entirely.  logLevel: process.env.NODE_ENV === "production" ? "warn" : "info",   // Off by default so it cannot leak a live process view.  dashboard: process.env.NODE_ENV !== "production",   // In adaptive mode this is a multiplier of each query's own rolling p95.  slowQueryThresholdMs: 100,  slowQueryMode: "fixed",});
Warning
Create the observer once. Two instances means two async contexts, and a span started in one is never a parent of a span started in the other. Unknown config keys throw at construction rather than being ignored, so a typo like runtimeMetric fails at boot instead of silently leaving the feature off.

It reports what it does not know

An analysis that gives a confident answer when it is guessing is worse than no analysis, because it gets trusted. Every conclusion here carries the evidence behind it and the reasons it might be wrong, and gaps are reported rather than smoothed over.

tsts
const a = observer.analyzeIncident(fingerprint)!; a.suspects[0]?.label;      // "Latency concentrated in db.query"a.suspects[0]?.confidence; // "likely" | "possible" | "unknown"a.suspects[0]?.score;      // 0..1 — ranks suspects, not a probabilitya.suspects[0]?.evidence;   // what supports ita.suspects[0]?.caveats;    // why it could still be wronga.unknowns;                // never empty — see below

unknowns always carries a standing disclosure that the correlation is heuristic and based only on in-memory telemetry. Even a well-formed incident with plenty of data gets that line, so a consumer cannot render a bare “Likely cause” with nothing beside it.

Where the data cannot support a claim, the engine declines rather than guessing:

Refusals are explicit, not silentts
// Fewer than 20 samples for a baseline{ anomalous: false, confidence: "unknown",  caveats: ["Only 4 sample(s) of dbLatencyMs…; a baseline needs 20."] } // Metric with zero observed variance — no noise scale to score against{ z: null, anomalous: false,  caveats: ["…has near-zero variance in this window…"] } // Zero baseline for a ratio{ changePct: null }        // not Infinity // SLO with too little traffic{ status: "no-data", achieved: null } // First release, with nothing to compare against{ status: "unknown" }      // not "healthy" // Memory growing, but the fit is weak{ verdict: "stable", caveats: ["Trend fit is weak (R²=0.21)…"] }

Issues

Errors are grouped into issues. Each issue holds many occurrences — the individual events — and every occurrence links back to its trace.

tsts
ErrorIssue├── fingerprint    groups every occurrence of the same failure├── status         "unresolved" | "resolved" | "ignored"├── firstSeen / lastSeen├── count          total occurrences├── name, message  of the first occurrence├── statusCode     when the error carried one├── topFrame       the in-app frame, for jumping to source├── routes[]       every route it was seen from├── sparkline[]    5-minute buckets for trend└── occurrences[]  most recent 25

Triage

tsts
observer.errors.applyAction(fingerprint, "resolve");observer.errors.applyAction(fingerprint, "ignore");observer.errors.applyAction(fingerprint, "unresolve"); observer.errors.stats();   // { unresolved, events, … } // Analysing through the reporter rather than the raw map keeps grouping,// counting and the event log consistent with each other.observer.recordError(preBuiltOccurrence);
Note
recordError() groups into the issue queue. An earlier version only pushed onto the event log, so errors arriving as pre-built occurrences never reached the dashboard’s Errors view — which is exactly what that view reads.

Grouping

The fingerprint comes from the error name, a normalized message and the most salient in-application stack frames. Normalization is what stops one bug becoming a thousand issues:

tsts
User 123 failed   ─┐User 456 failed   ─┼─▶  one fingerprint → one issueUser 789 failed   ─┘ computeSmartFingerprint(error, parseStackTrace(error.stack))

Tracing

tsts
const span = observer.tracer.startSpan("checkout.service", {  kind: "internal",  attributes: { "cart.items": 12 },}); try {  const receipt = await processPayment(cart);  span.ok();  return receipt;} catch (err) {  span.recordError(err);  throw err;                       // observe, never swallow} finally {  observer.tracer.endSpan(span);} // Or let the tracer handle it:await observer.traceDb("orders.findMany", () => db.orders.findMany({ … }));await observer.traceJob("send-receipt", () => send(receipt));

instrument() wraps a fetch handler so every request becomes a server span with a trace id in the response headers:

tsts
const instrumented = observer.instrument(  async (req) => handle(req),  (req) => new URL(req.url).pathname,); Bun.serve({ fetch: (req, server) => instrumented(req) });
Warning
A handler that throws still records its status. Only recording it on the success path left http.status_code unset, so every consumer defaulted it to 200 — a crashed endpoint was counted as a fast successful request, inflating Apdex and hiding the failure in the transaction log.

Metrics and adaptive baselines

The observer retains a bounded sample history. Everything trend-shaped — adaptive thresholds, memory growth, backlog direction — is derived from it rather than from a current value.

Per-query thresholdsts
createObserver({  slowQueryThresholdMs: 3,        // multiplier, not milliseconds  slowQueryMode: "adaptive",  baselineWindowMs: 15 * 60_000,}); // A table that normally takes 400ms is no longer all flagged;// a table that normally takes 5ms still gets caught.observer.detectNPlusOne();       // repeated identical queriesobserver.detectAnomalies({ cpuPercent: 88, queueDepth: 900 });
Note
During warm-up, adaptive queries are recorded but left unclassified. Treating the multiplier as a millisecond budget flagged nearly everything in the first few dozen calls, which is worse than saying “not enough data yet”. Baselines use median and MAD rather than mean and standard deviation, because a single stall drags a mean upward and inflates a σ — hiding the event worth reporting.

Memory trend

tsts
const t = observer.memoryTrend("rssMb");t.growthMbPerHour;   // least-squares slopet.rSquared;          // fit qualityt.verdict;           // "growing" | "stable" | "shrinking" | "insufficient-data" t.caveats;  // sustained growth is consistent with a leak, never proof of one

Releases

Set release and the observer records its own build at startup, from its own config rather than from a caller — so the timeline cannot claim a version that isn’t the one running.

tsts
observer.releaseHealth();// [{ release, requests, errors, errorRatePct, p50Ms, p95Ms, p99Ms,//    slowQueries, newErrors[], affectedRoutes[], status }] observer.whatChanged();// { release, previousRelease, deltas[], newErrors[], fixedErrors[], summary[] } // status is "unknown" for the first release and for any release whose// comparison window has fewer than 20 requests — "healthy" would be an// assumption dressed as a measurement.
Warning
“Fixed” errors key off lastSeen, not firstSeen. Keying off firstSeen reported every long-standing error as fixed on each release simply because it predates the deploy — including ones still firing hundreds of times a minute.

Incident correlation

tsts
const a = observer.analyzeIncident(fingerprint)!; a.timeline;            // ordered request/span/error events with offsetsa.suspects;            // ranked, each with evidence and caveatsa.suggestedActions;    // each with a rationale and a risk levela.unknowns;            // what the data could not answera.stats.traceReusePct; // does one trace explain most failures? observer.analyzeAllIncidents();   // every non-ignored issue, newest first

It correlates the failing trace, the transaction, slow queries on the same trace, release proximity, and rule-outs. A suspect that the evidence rules out is reported as a suspect:

tsts
{  id: "not-current-release",  label: "Not caused by the current release (1.8.4)",  confidence: "likely",  evidence: [{ kind: "deploy", summary: "Issue first seen … under 1.8.3" }],  caveats: ["A shared dependency or database change can affect every release             at once, so an older release does not rule out an external cause."],}

Everything it derives

tsts
observer.analyzeIncident(fp)      // timeline, suspects, actionsobserver.analyzeAllIncidents()observer.releaseHealth()observer.whatChanged(release?)observer.serviceMap()                // graph inferred from spansobserver.detectNPlusOne()observer.detectAnomalies(values)observer.memoryTrend(metric?)observer.evaluateSlos()observer.evaluateSlo(slo)observer.jobHealth()observer.incidentReport(fp)           // self-contained JSONobserver.saveGoldenTrace(name, traceId)observer.compareToGolden(name, traceId)observer.previewAction(request)observer.performAction(request, executor, opts)

N+1 detection

Database spans are grouped by the parent they were issued from. Repetition under one parent is the N+1 shape specifically — the same statement once per request is just traffic.

tsts
observer.detectNPlusOne();// [{ traceId, parentName, normalizedSql, calls: 101,//    eachMs: 5, avoidableMs: 500, severity: "high",//    caveats: ["Repetition alone does not prove a loop…"] }] // avoidableMs excludes the first call, which is necessary.

Dependency map

tsts
const map = observer.serviceMap();// nodes: api, database, cache, jobs, mail, realtime… map.limitations;// [//   "Nodes are inferred from span names and kinds, so this shows instrumented//    call paths rather than a declared architecture.",//   "No producer or consumer spans were recorded, so asynchronous edges…//    are missing.",//   "No outbound HTTP client spans were recorded…"// ]

Nodes are classified from span kinds and names. A route span is the incoming edge of the API regardless of which span kind the instrumentation assigned it — splitting a route on / to guess a service name turned POST /api/checkout into a service called post.

SLOs and error budgets

tsts
createObserver({  slos: [    { name: "Checkout availability", target: 99.9, windowMs: 30 * 86_400_000 },    { name: "Orders POST", target: 99, windowMs: 86_400_000,      query: { route: "/api/orders", method: "POST" } },  ],}); observer.evaluateSlos();// [{ name, target, requests, achieved, errorBudgetRemaining, burnRate,//    status: "healthy" | "at-risk" | "breached" | "no-data", caveats[] }]
Note
A 5xx counts against availability; a 4xx does not, because it is a served answer rather than an outage. Burn rate is measured against the interval the traffic actually covers, not the configured window — deriving it from the window start made a 30-second burst in a 30-day window look like full coverage and understated burn rate by orders of magnitude.

Job health

tsts
observer.jobHealth();// { queued, running, delayed, completed, failed, dead,//   successRate, backlog, growthPerMinute, caveats[] } // backlog is "unknown" until there are 20 depth samples, and a small drift// against a large queue is "stable" — +1/s at 10k queued is noise.

Guarded actions

Destructive operations are previewed before they run, require confirmation, and require an idempotency key so a double-click cannot purge twice.

tsts
const preview = observer.previewAction({ kind: "purge-dead-jobs" });preview.risk;                 // "safe" | "caution" | "destructive"preview.requiresConfirmation;  // truepreview.target;                // "all matching entries"preview.reason;                // "Permanently discards failed jobs. They                               //  cannot be replayed afterwards." const outcome = observer.performAction(  { kind: "purge-dead-jobs", idempotencyKey: requestId },  execute,  { confirmed: true },); outcome.deduplicated;  // true if the same key was already performed
Warning
Blocked attempts are audited too — “someone tried to purge the queue at 3am” is itself the record worth keeping. The idempotency ledger lives on the observer instance; building it per request emptied it every time and a retried destructive action ran twice.

Golden traces

tsts
observer.saveGoldenTrace("Cart — healthy", traceId);observer.compareToGolden("Cart — healthy", traceId);// { verdict, differences[] } // differences: { step, kind: "extra" | "missing" | "slower" | "faster",//               deltaMs, ratio, note } // A missing step's note says it may be an improvement or a removed step —// reporting it as a regression would train people to ignore it.

A baseline records per-step timings, not just step names. Without them the comparison can detect a new or missing step but never a slowdown, which is the main reason to keep a golden trace at all.

Trace to code

tsts
import { FileSystemTraceToCode } from "yatta/observe";import fs from "node:fs"; // Opt-in: production builds usually ship no sources, and reading files off a// running server is not something to enable implicitly.observer.sourceResolver = new FileSystemTraceToCode(  (p) => fs.existsSync(p) ? fs.readFileSync(p, "utf-8") : null,  process.cwd(),); observer.sourceFor({ filePath, line, column });// { lines: [{ number, text, isTarget }], missing?, reason?, caveats[] }
Warning
The resolver refuses any path outside the project root. A stack frame can carry any path an attacker put in an error message, so an unguarded read would turn a crafted trace into an arbitrary file read.

Incident reports

incidentReport() returns a self-contained JSON payload designed for a language model or a colleague with no access to the process. Conclusions travel with their evidence and their weaknesses.

tsts
const report = observer.incidentReport(fingerprint);// {//   schemaVersion: 1,//   service: { name, release, version, environment },//   summary: { errorName, message, occurrences, affectedRoutes, … },//   likelyCauses: [{ claim, confidence, supportingEvidence[],//                    reasonsItCouldBeWrong[] }],//   knownUnknowns: [],//   timeline: [],//   suggestedNextSteps: [{ action, why }],//   collectionGaps: []// }
Note
confidence is rendered as a string carrying the score —“possible (score 0.4)” rather than a bare possible — so the number cannot be dropped in transit.

HTTP endpoints

tsts
GET  /_yatta                              dashboardGET  /_yatta/api/stream                      SSE event streamGET  /_yatta/api/telemetry                   full application stateGET  /_yatta/api/traces/:id                   waterfallGET  /_yatta/api/errors                      issue queueGET  /_yatta/api/logs POST /_yatta/api/errors/:action              resolve | ignore | unresolvePOST /_yatta/api/jobs/replay-deadPOST /_yatta/api/jobs/purge-deadPOST /_yatta/api/cache/clear GET  /_yatta/api/analysis/state              every analysis, one callGET  /_yatta/api/analysis/incident/:fpGET  /_yatta/api/analysis/report/:fpGET  /_yatta/api/analysis/goldensPOST /_yatta/api/analysis/goldens/:name?trace=GET  /_yatta/api/analysis/compare/:name?trace=POST /_yatta/api/analysis/action/:kind
Warning
Actions route through the guard rather than executing inline, so confirmation and idempotency cannot be bypassed by calling the endpoint directly — which is exactly what someone would do.

What it cannot see

Stated by the features themselves rather than left to be discovered.

  • History is in-memory only. A restart loses it. Retention is bounded (720 samples by default) precisely so the observer cannot become the leak it exists to diagnose.
  • No outbound HTTP instrumentation. serviceMap() reports this as a limitation; external dependencies appear only if a call site was wrapped explicitly.
  • Source navigation needs a resolver, and production bundles ship no sources.
  • SLO windows are bounded by retention. A 30-day objective measured against an in-process buffer will undercount, and says so in caveats.