Telemetry

The 5-surface model#

Telemetry has two axes. The user/host axis measures whole-conversation usage (what the user spent); the developer/surface axis measures what the connector costs — now across all five developer surfaces.

The two axes#

User / host axis

Whole-conversation host usage — what the USER spent across the entire conversation.

The model_turn host-native hook (live, exact) + the host-scan CLI-log readers. Surfaced by the Host/User + Host-native-turns leaderboards.

Developer / surface axis

What the CONNECTOR costs — the footprint your connector imposes, now across ALL FIVE surfaces.

server + hooks measured live (RUNTIME store rows); commands + skills + subagents computed on demand as STATIC context footprints.

One table, four vocabularies#

The telemetry types use four names for closely-related things — the developer surface, its EventScope(s), its SurfaceKind, and whether it is RUNTIME-measured or a STATIC footprint. On the developer surfaces EventScope and SurfaceKind co-vary, so this one table lines all four up. The detailed per-vocabulary tables follow below.

SurfaceEventScope(s)SurfaceKindRUNTIME / STATICWhat it is
servercall, tool_defsserver
RUNTIME
Serve-proxy per-tool round-trips + the one-time tools/list schema overhead.
hookshookhook
RUNTIME
One row per hook dispatch through the home-bin entrypoint.
command— (no store row)command
STATIC
Tokenized context footprint computed on demand; never a usage row.
skill— (no store row)skill
STATIC
Tokenized context footprint computed on demand; never a usage row.
subagent— (no store row)subagent
STATIC
Tokenized context footprint computed on demand; never a usage row.
(not a developer surface)model_turn— (none)—A host-native whole-conversation turn (e.g. Gemini/Antigravity AfterModel). Excluded from the per-MCP and per-surface views — it has its own leaderboard board.

The five developer surfaces#

Two surfaces are RUNTIME (measured live, producing store rows): server (per-MCP-tool call + tool_defs via the serve-proxy) and hooks (per-event, measured at the home-bin hook entrypoint). Three are STATIC footprints computed on-demand from the connector — command, skill, subagent — the context cost the host pays to load them. Static footprints are sizes, not usage, and are never written as fake rows.

SurfaceKindWhat is measuredDetail
server
RUNTIME
per-MCP-tool call + tool_defsThe serve-proxy tokenizes each tools/call round-trip (scope call) and the one-time tools/list schema overhead (scope tool_defs). surfaceKind "server" (the backward-compatible default for legacy rows).
hooks
RUNTIME
per-event hook dispatchMeasured at the home-bin hook entrypoint (src/runtime/hook-entrypoint): one row per RUNTIME hook dispatch (scope "hook", surfaceKind "hook"). Input = the inbound normalized event payload; output = the HookResponse that becomes context/decision. The per-item name IS the event (e.g. SessionStart). Fail-open: a telemetry error never breaks the hook.
command
STATIC
context footprint (description + prompt + argumentHint)Computed on demand from the connector (surface-footprint.ts) — a tokenized footprint of the context the host loads, NOT runtime usage. Never written as a store row.
skill
STATIC
context footprint (description + body + resources)Static footprint of SKILL.md + every resource value (sorted by path for determinism). Computed on demand, never a usage row.
subagent
STATIC
context footprint (description + prompt)Static footprint of the subagent's description + system prompt. Computed on demand, never a usage row.

hook scope + surfaceKind are new

The runtime hook surface adds a new EventScope value "hook" and stamps surfaceKind: "hook" on each row. Measurement happens at the home-bin hook entrypoint and is fail-open: a telemetry error can never break a host's hook.

EventScope & SurfaceKind#

Every store row carries an EventScope (what it measures) and an optional SurfaceKind (which developer surface). The four scopes are distinct origins that must never be summed:

EventScopeMeaning
"call"One per-MCP tools/call round-trip (serve-proxy bytes). The headline per-tool cost.
"tool_defs"The one-time tools/list schema overhead (serve-proxy). Counted as tokens but never as a call.
"model_turn"A WHOLE-CONVERSATION host-native turn the host reported (e.g. Gemini/Antigravity AfterModel usageMetadata). EXCLUDED from the per-MCP/per-surface views — its own leaderboard section.
"hook"One RUNTIME hook dispatch through the home-bin entrypoint. The developer-axis hook surface — measured live, like call.
SurfaceKindMeaning
"server"RUNTIME serve-proxy rows (call / tool_defs). The backward-compatible default: rows written before surfaceKind existed read as server.
"hook"RUNTIME hook-entrypoint rows (scope hook), stamped explicitly.
"command"STATIC command footprint — never a store row.
"skill"STATIC skill footprint — never a store row.
"subagent"STATIC subagent footprint — never a store row.

surfaceKind is optional and backward-compatible: rows written before the field existed (every legacy serve-proxy call/tool_defs row) lack it and are read as server. The command/skill/subagent kinds only ever appear on static footprints — they never produce store rows.

Local-first, zero-egress, opt-out#

  • Local-first. Everything is tokenized locally and stored under the home data-root — aggregate counts only, never raw arguments or results.
  • Zero network egress by default. The hot path makes no network call; only the opt-in calibration sampler ever sends content off-box.
  • Opt-out. AGENT_CONNECTOR_TELEMETRY=0 is a global kill switch honored by both the serve-proxy and the hook runtime. telemetry: { enabled: false } suppresses the serve-proxy telemetry wrap at install time (shouldWrapForTelemetry gates on enabled === true), but is NOT currently consulted by the hook runtime — an installed hook still records telemetry rows unless the env var is set.

Confidence sources#

Every row (and every static footprint) carries one confidence source so an estimate is never read as exact — see the confidence sources table. Static footprints are labeled with the tokenizer source for the connector's family (tokenizer-exact for OpenAI-family, tokenizer-approx otherwise).

The per-surface leaderboard#

agent-connector telemetry leaderboard --by mcp|tool|surface ranks the per-MCP telemetry by connector (the default --by mcp, "which MCP server costs the most"), by tool, or — new — by developer-axis surface. The --by surface view folds the runtime server/hook store rows together with the static command/skill/subagent footprints of the registered connector(s). Its columns:

ColumnMeaning
SURFACEThe surfaceKind (server | hook | command | skill | subagent).
NAMEThe per-item name: the tool name (server), the event name (hook), or the command/skill/subagent name (static).
INInput tokens. For static rows, the whole footprint sits here.
OUTOutput tokens. Always 0 for static rows.
TOTALIN + OUT for the (surface, name) group.
KINDruntime (live usage, aggregated from the store) vs static (a context-load footprint). The distinction is never silently conflated.
terminal
text
$ agent-connector telemetry leaderboard --by surface

SURFACE   NAME              IN      OUT    TOTAL  KIND
----------------------------------------------------------
server    query           4,210   9,880  14,090  runtime
hook      PreToolUse      1,120     640   1,760  runtime
skill     deep-research     980       0     980  static
command   deploy            612       0     612  static
subagent  reviewer          540       0     540  static
----------------------------------------------------------
TOTAL                     7,462  10,520  17,982

note: KIND=static rows are the tokenized FOOTPRINT a command/skill/subagent
imposes on a host that loads it as context — not intercepted usage rows.

Sizes are never summed with usage

Static footprints are sizes (the context-load cost of a surface), not runtime usage. The KIND column keeps runtime vs static explicit so the two are never silently conflated, and the whole-conversation model_turn rows are excluded from this view entirely (they get their own leaderboard section).