Observability
Tracing, metrics, logs, issue grouping, incident correlation, release intelligence, adaptive baselines, SLOs, N+1 detection and golden traces — in-process.
yatta/observe is built into the framework. One Observer records spans, metrics, logs and issues, and derives analysis from them. Spans propagate through AsyncLocalStorage, so a trace follows a request into a job worker and out over SSE without passing anything around. Nothing leaves the process.
Setup
import { createObserver } from "yatta/observe"; export const observer = createObserver({ service: "app", release: process.env.GIT_SHA, // Logical OR, not ?? — an empty NODE_ENV is common in Docker and CI, and a // nullish coalesce treats "" as set, so a production container would read as // development and serve the dashboard. environment: process.env.NODE_ENV || "development", bufferSize: 2000, // Everything below this level is dropped before it is masked, buffered, // breadcrumb'd or broadcast. “silent” turns logging off entirely. logLevel: process.env.NODE_ENV === "production" ? "warn" : "info", // Off by default so it cannot leak a live process view. dashboard: process.env.NODE_ENV !== "production", // In adaptive mode this is a multiplier of each query's own rolling p95. slowQueryThresholdMs: 100, slowQueryMode: "fixed",});runtimeMetric fails at boot instead of silently leaving the feature off.It reports what it does not know
An analysis that gives a confident answer when it is guessing is worse than no analysis, because it gets trusted. Every conclusion here carries the evidence behind it and the reasons it might be wrong, and gaps are reported rather than smoothed over.
const a = observer.analyzeIncident(fingerprint)!; a.suspects[0]?.label; // "Latency concentrated in db.query"a.suspects[0]?.confidence; // "likely" | "possible" | "unknown"a.suspects[0]?.score; // 0..1 — ranks suspects, not a probabilitya.suspects[0]?.evidence; // what supports ita.suspects[0]?.caveats; // why it could still be wronga.unknowns; // never empty — see belowunknowns always carries a standing disclosure that the correlation is heuristic and based only on in-memory telemetry. Even a well-formed incident with plenty of data gets that line, so a consumer cannot render a bare “Likely cause” with nothing beside it.
Where the data cannot support a claim, the engine declines rather than guessing:
// Fewer than 20 samples for a baseline{ anomalous: false, confidence: "unknown", caveats: ["Only 4 sample(s) of dbLatencyMs…; a baseline needs 20."] } // Metric with zero observed variance — no noise scale to score against{ z: null, anomalous: false, caveats: ["…has near-zero variance in this window…"] } // Zero baseline for a ratio{ changePct: null } // not Infinity // SLO with too little traffic{ status: "no-data", achieved: null } // First release, with nothing to compare against{ status: "unknown" } // not "healthy" // Memory growing, but the fit is weak{ verdict: "stable", caveats: ["Trend fit is weak (R²=0.21)…"] }Issues
Errors are grouped into issues. Each issue holds many occurrences — the individual events — and every occurrence links back to its trace.
ErrorIssue├── fingerprint groups every occurrence of the same failure├── status "unresolved" | "resolved" | "ignored"├── firstSeen / lastSeen├── count total occurrences├── name, message of the first occurrence├── statusCode when the error carried one├── topFrame the in-app frame, for jumping to source├── routes[] every route it was seen from├── sparkline[] 5-minute buckets for trend└── occurrences[] most recent 25Triage
observer.errors.applyAction(fingerprint, "resolve");observer.errors.applyAction(fingerprint, "ignore");observer.errors.applyAction(fingerprint, "unresolve"); observer.errors.stats(); // { unresolved, events, … } // Analysing through the reporter rather than the raw map keeps grouping,// counting and the event log consistent with each other.observer.recordError(preBuiltOccurrence);recordError() groups into the issue queue. An earlier version only pushed onto the event log, so errors arriving as pre-built occurrences never reached the dashboard’s Errors view — which is exactly what that view reads.Grouping
The fingerprint comes from the error name, a normalized message and the most salient in-application stack frames. Normalization is what stops one bug becoming a thousand issues:
User 123 failed ─┐User 456 failed ─┼─▶ one fingerprint → one issueUser 789 failed ─┘ computeSmartFingerprint(error, parseStackTrace(error.stack))Tracing
const span = observer.tracer.startSpan("checkout.service", { kind: "internal", attributes: { "cart.items": 12 },}); try { const receipt = await processPayment(cart); span.ok(); return receipt;} catch (err) { span.recordError(err); throw err; // observe, never swallow} finally { observer.tracer.endSpan(span);} // Or let the tracer handle it:await observer.traceDb("orders.findMany", () => db.orders.findMany({ … }));await observer.traceJob("send-receipt", () => send(receipt));instrument() wraps a fetch handler so every request becomes a server span with a trace id in the response headers:
const instrumented = observer.instrument( async (req) => handle(req), (req) => new URL(req.url).pathname,); Bun.serve({ fetch: (req, server) => instrumented(req) });http.status_code unset, so every consumer defaulted it to 200 — a crashed endpoint was counted as a fast successful request, inflating Apdex and hiding the failure in the transaction log.Metrics and adaptive baselines
The observer retains a bounded sample history. Everything trend-shaped — adaptive thresholds, memory growth, backlog direction — is derived from it rather than from a current value.
createObserver({ slowQueryThresholdMs: 3, // multiplier, not milliseconds slowQueryMode: "adaptive", baselineWindowMs: 15 * 60_000,}); // A table that normally takes 400ms is no longer all flagged;// a table that normally takes 5ms still gets caught.observer.detectNPlusOne(); // repeated identical queriesobserver.detectAnomalies({ cpuPercent: 88, queueDepth: 900 });Memory trend
const t = observer.memoryTrend("rssMb");t.growthMbPerHour; // least-squares slopet.rSquared; // fit qualityt.verdict; // "growing" | "stable" | "shrinking" | "insufficient-data" t.caveats; // sustained growth is consistent with a leak, never proof of oneReleases
Set release and the observer records its own build at startup, from its own config rather than from a caller — so the timeline cannot claim a version that isn’t the one running.
observer.releaseHealth();// [{ release, requests, errors, errorRatePct, p50Ms, p95Ms, p99Ms,// slowQueries, newErrors[], affectedRoutes[], status }] observer.whatChanged();// { release, previousRelease, deltas[], newErrors[], fixedErrors[], summary[] } // status is "unknown" for the first release and for any release whose// comparison window has fewer than 20 requests — "healthy" would be an// assumption dressed as a measurement.lastSeen, not firstSeen. Keying off firstSeen reported every long-standing error as fixed on each release simply because it predates the deploy — including ones still firing hundreds of times a minute.Incident correlation
const a = observer.analyzeIncident(fingerprint)!; a.timeline; // ordered request/span/error events with offsetsa.suspects; // ranked, each with evidence and caveatsa.suggestedActions; // each with a rationale and a risk levela.unknowns; // what the data could not answera.stats.traceReusePct; // does one trace explain most failures? observer.analyzeAllIncidents(); // every non-ignored issue, newest firstIt correlates the failing trace, the transaction, slow queries on the same trace, release proximity, and rule-outs. A suspect that the evidence rules out is reported as a suspect:
{ id: "not-current-release", label: "Not caused by the current release (1.8.4)", confidence: "likely", evidence: [{ kind: "deploy", summary: "Issue first seen … under 1.8.3" }], caveats: ["A shared dependency or database change can affect every release at once, so an older release does not rule out an external cause."],}Everything it derives
observer.analyzeIncident(fp) // timeline, suspects, actionsobserver.analyzeAllIncidents()observer.releaseHealth()observer.whatChanged(release?)observer.serviceMap() // graph inferred from spansobserver.detectNPlusOne()observer.detectAnomalies(values)observer.memoryTrend(metric?)observer.evaluateSlos()observer.evaluateSlo(slo)observer.jobHealth()observer.incidentReport(fp) // self-contained JSONobserver.saveGoldenTrace(name, traceId)observer.compareToGolden(name, traceId)observer.previewAction(request)observer.performAction(request, executor, opts)N+1 detection
Database spans are grouped by the parent they were issued from. Repetition under one parent is the N+1 shape specifically — the same statement once per request is just traffic.
observer.detectNPlusOne();// [{ traceId, parentName, normalizedSql, calls: 101,// eachMs: 5, avoidableMs: 500, severity: "high",// caveats: ["Repetition alone does not prove a loop…"] }] // avoidableMs excludes the first call, which is necessary.Dependency map
const map = observer.serviceMap();// nodes: api, database, cache, jobs, mail, realtime… map.limitations;// [// "Nodes are inferred from span names and kinds, so this shows instrumented// call paths rather than a declared architecture.",// "No producer or consumer spans were recorded, so asynchronous edges…// are missing.",// "No outbound HTTP client spans were recorded…"// ]Nodes are classified from span kinds and names. A route span is the incoming edge of the API regardless of which span kind the instrumentation assigned it — splitting a route on / to guess a service name turned POST /api/checkout into a service called post.
SLOs and error budgets
createObserver({ slos: [ { name: "Checkout availability", target: 99.9, windowMs: 30 * 86_400_000 }, { name: "Orders POST", target: 99, windowMs: 86_400_000, query: { route: "/api/orders", method: "POST" } }, ],}); observer.evaluateSlos();// [{ name, target, requests, achieved, errorBudgetRemaining, burnRate,// status: "healthy" | "at-risk" | "breached" | "no-data", caveats[] }]Job health
observer.jobHealth();// { queued, running, delayed, completed, failed, dead,// successRate, backlog, growthPerMinute, caveats[] } // backlog is "unknown" until there are 20 depth samples, and a small drift// against a large queue is "stable" — +1/s at 10k queued is noise.Guarded actions
Destructive operations are previewed before they run, require confirmation, and require an idempotency key so a double-click cannot purge twice.
const preview = observer.previewAction({ kind: "purge-dead-jobs" });preview.risk; // "safe" | "caution" | "destructive"preview.requiresConfirmation; // truepreview.target; // "all matching entries"preview.reason; // "Permanently discards failed jobs. They // cannot be replayed afterwards." const outcome = observer.performAction( { kind: "purge-dead-jobs", idempotencyKey: requestId }, execute, { confirmed: true },); outcome.deduplicated; // true if the same key was already performedGolden traces
observer.saveGoldenTrace("Cart — healthy", traceId);observer.compareToGolden("Cart — healthy", traceId);// { verdict, differences[] } // differences: { step, kind: "extra" | "missing" | "slower" | "faster",// deltaMs, ratio, note } // A missing step's note says it may be an improvement or a removed step —// reporting it as a regression would train people to ignore it.A baseline records per-step timings, not just step names. Without them the comparison can detect a new or missing step but never a slowdown, which is the main reason to keep a golden trace at all.
Trace to code
import { FileSystemTraceToCode } from "yatta/observe";import fs from "node:fs"; // Opt-in: production builds usually ship no sources, and reading files off a// running server is not something to enable implicitly.observer.sourceResolver = new FileSystemTraceToCode( (p) => fs.existsSync(p) ? fs.readFileSync(p, "utf-8") : null, process.cwd(),); observer.sourceFor({ filePath, line, column });// { lines: [{ number, text, isTarget }], missing?, reason?, caveats[] }Incident reports
incidentReport() returns a self-contained JSON payload designed for a language model or a colleague with no access to the process. Conclusions travel with their evidence and their weaknesses.
const report = observer.incidentReport(fingerprint);// {// schemaVersion: 1,// service: { name, release, version, environment },// summary: { errorName, message, occurrences, affectedRoutes, … },// likelyCauses: [{ claim, confidence, supportingEvidence[],// reasonsItCouldBeWrong[] }],// knownUnknowns: [],// timeline: [],// suggestedNextSteps: [{ action, why }],// collectionGaps: []// }confidence is rendered as a string carrying the score —“possible (score 0.4)” rather than a bare possible — so the number cannot be dropped in transit.HTTP endpoints
GET /_yatta dashboardGET /_yatta/api/stream SSE event streamGET /_yatta/api/telemetry full application stateGET /_yatta/api/traces/:id waterfallGET /_yatta/api/errors issue queueGET /_yatta/api/logs POST /_yatta/api/errors/:action resolve | ignore | unresolvePOST /_yatta/api/jobs/replay-deadPOST /_yatta/api/jobs/purge-deadPOST /_yatta/api/cache/clear GET /_yatta/api/analysis/state every analysis, one callGET /_yatta/api/analysis/incident/:fpGET /_yatta/api/analysis/report/:fpGET /_yatta/api/analysis/goldensPOST /_yatta/api/analysis/goldens/:name?trace=GET /_yatta/api/analysis/compare/:name?trace=POST /_yatta/api/analysis/action/:kindWhat it cannot see
Stated by the features themselves rather than left to be discovered.
- History is in-memory only. A restart loses it. Retention is bounded (720 samples by default) precisely so the observer cannot become the leak it exists to diagnose.
- No outbound HTTP instrumentation.
serviceMap()reports this as a limitation; external dependencies appear only if a call site was wrapped explicitly. - Source navigation needs a resolver, and production bundles ship no sources.
- SLO windows are bounded by retention. A 30-day objective measured against an in-process buffer will undercount, and says so in
caveats.