Telemetry
The 5-surface model#
Telemetry has two axes. The user/host axis measures whole-conversation usage (what the user spent); the developer/surface axis measures what the connector costs — now across all five developer surfaces.
The two axes#
Whole-conversation host usage — what the USER spent across the entire conversation.
The model_turn host-native hook (live, exact) + the host-scan CLI-log readers. Surfaced by the Host/User + Host-native-turns leaderboards.
What the CONNECTOR costs — the footprint your connector imposes, now across ALL FIVE surfaces.
server + hooks measured live (RUNTIME store rows); commands + skills + subagents computed on demand as STATIC context footprints.
One table, four vocabularies#
The telemetry types use four names for closely-related things — the developer surface, its EventScope(s), its SurfaceKind, and whether it is RUNTIME-measured or a STATIC footprint. On the developer surfaces EventScope and SurfaceKind co-vary, so this one table lines all four up. The detailed per-vocabulary tables follow below.
| Surface | EventScope(s) | SurfaceKind | RUNTIME / STATIC | What it is |
|---|---|---|---|---|
server | call, tool_defs | server | RUNTIME | Serve-proxy per-tool round-trips + the one-time tools/list schema overhead. |
hooks | hook | hook | RUNTIME | One row per hook dispatch through the home-bin entrypoint. |
command | — (no store row) | command | STATIC | Tokenized context footprint computed on demand; never a usage row. |
skill | — (no store row) | skill | STATIC | Tokenized context footprint computed on demand; never a usage row. |
subagent | — (no store row) | subagent | STATIC | Tokenized context footprint computed on demand; never a usage row. |
(not a developer surface) | model_turn | — (none) | — | A host-native whole-conversation turn (e.g. Gemini/Antigravity AfterModel). Excluded from the per-MCP and per-surface views — it has its own leaderboard board. |
The five developer surfaces#
Two surfaces are RUNTIME (measured live, producing store rows): server (per-MCP-tool call + tool_defs via the serve-proxy) and hooks (per-event, measured at the home-bin hook entrypoint). Three are STATIC footprints computed on-demand from the connector — command, skill, subagent — the context cost the host pays to load them. Static footprints are sizes, not usage, and are never written as fake rows.
| Surface | Kind | What is measured | Detail |
|---|---|---|---|
server | RUNTIME | per-MCP-tool call + tool_defs | The serve-proxy tokenizes each tools/call round-trip (scope call) and the one-time tools/list schema overhead (scope tool_defs). surfaceKind "server" (the backward-compatible default for legacy rows). |
hooks | RUNTIME | per-event hook dispatch | Measured at the home-bin hook entrypoint (src/runtime/hook-entrypoint): one row per RUNTIME hook dispatch (scope "hook", surfaceKind "hook"). Input = the inbound normalized event payload; output = the HookResponse that becomes context/decision. The per-item name IS the event (e.g. SessionStart). Fail-open: a telemetry error never breaks the hook. |
command | STATIC | context footprint (description + prompt + argumentHint) | Computed on demand from the connector (surface-footprint.ts) — a tokenized footprint of the context the host loads, NOT runtime usage. Never written as a store row. |
skill | STATIC | context footprint (description + body + resources) | Static footprint of SKILL.md + every resource value (sorted by path for determinism). Computed on demand, never a usage row. |
subagent | STATIC | context footprint (description + prompt) | Static footprint of the subagent's description + system prompt. Computed on demand, never a usage row. |
hook scope + surfaceKind are new
The runtime hook surface adds a newEventScope value "hook" and stamps surfaceKind: "hook" on each row. Measurement happens at the home-bin hook entrypoint and is fail-open: a telemetry error can never break a host's hook.EventScope & SurfaceKind#
Every store row carries an EventScope (what it measures) and an optional SurfaceKind (which developer surface). The four scopes are distinct origins that must never be summed:
| EventScope | Meaning |
|---|---|
"call" | One per-MCP tools/call round-trip (serve-proxy bytes). The headline per-tool cost. |
"tool_defs" | The one-time tools/list schema overhead (serve-proxy). Counted as tokens but never as a call. |
"model_turn" | A WHOLE-CONVERSATION host-native turn the host reported (e.g. Gemini/Antigravity AfterModel usageMetadata). EXCLUDED from the per-MCP/per-surface views — its own leaderboard section. |
"hook" | One RUNTIME hook dispatch through the home-bin entrypoint. The developer-axis hook surface — measured live, like call. |
| SurfaceKind | Meaning |
|---|---|
"server" | RUNTIME serve-proxy rows (call / tool_defs). The backward-compatible default: rows written before surfaceKind existed read as server. |
"hook" | RUNTIME hook-entrypoint rows (scope hook), stamped explicitly. |
"command" | STATIC command footprint — never a store row. |
"skill" | STATIC skill footprint — never a store row. |
"subagent" | STATIC subagent footprint — never a store row. |
surfaceKind is optional and backward-compatible: rows written before the field existed (every legacy serve-proxy call/tool_defs row) lack it and are read as server. The command/skill/subagent kinds only ever appear on static footprints — they never produce store rows.
Local-first, zero-egress, opt-out#
- Local-first. Everything is tokenized locally and stored under the home data-root — aggregate counts only, never raw arguments or results.
- Zero network egress by default. The hot path makes no network call; only the opt-in calibration sampler ever sends content off-box.
- Opt-out.
AGENT_CONNECTOR_TELEMETRY=0is a global kill switch honored by both the serve-proxy and the hook runtime.telemetry: { enabled: false }suppresses the serve-proxy telemetry wrap at install time (shouldWrapForTelemetrygates onenabled === true), but is NOT currently consulted by the hook runtime — an installed hook still records telemetry rows unless the env var is set.
Confidence sources#
Every row (and every static footprint) carries one confidence source so an estimate is never read as exact — see the confidence sources table. Static footprints are labeled with the tokenizer source for the connector's family (tokenizer-exact for OpenAI-family, tokenizer-approx otherwise).
The per-surface leaderboard#
agent-connector telemetry leaderboard --by mcp|tool|surface ranks the per-MCP telemetry by connector (the default --by mcp, "which MCP server costs the most"), by tool, or — new — by developer-axis surface. The --by surface view folds the runtime server/hook store rows together with the static command/skill/subagent footprints of the registered connector(s). Its columns:
| Column | Meaning |
|---|---|
SURFACE | The surfaceKind (server | hook | command | skill | subagent). |
NAME | The per-item name: the tool name (server), the event name (hook), or the command/skill/subagent name (static). |
IN | Input tokens. For static rows, the whole footprint sits here. |
OUT | Output tokens. Always 0 for static rows. |
TOTAL | IN + OUT for the (surface, name) group. |
KIND | runtime (live usage, aggregated from the store) vs static (a context-load footprint). The distinction is never silently conflated. |
$ agent-connector telemetry leaderboard --by surface
SURFACE NAME IN OUT TOTAL KIND
----------------------------------------------------------
server query 4,210 9,880 14,090 runtime
hook PreToolUse 1,120 640 1,760 runtime
skill deep-research 980 0 980 static
command deploy 612 0 612 static
subagent reviewer 540 0 540 static
----------------------------------------------------------
TOTAL 7,462 10,520 17,982
note: KIND=static rows are the tokenized FOOTPRINT a command/skill/subagent
imposes on a host that loads it as context — not intercepted usage rows.Sizes are never summed with usage
Static footprints are sizes (the context-load cost of a surface), not runtime usage. TheKIND column keeps runtime vs static explicit so the two are never silently conflated, and the whole-conversation model_turn rows are excluded from this view entirely (they get their own leaderboard section).