Pimp My IDE / Garage Dispatch
← Back to the garage
September 12, 2026 · agents / telemetry / review

Motion is not traction.

GitHub can now report active users, sessions, and messages from the dedicated VS Code Agents window. Useful. Also dangerously easy to turn into a productivity karaoke machine. Activity says the engine ran. It does not say the car went anywhere good.

The take: instrument adoption, but do not promote attendance into impact. Keep surface, traffic, delivered change, and production outcome in separate gears—and never let a rising message count impersonate engineering value.
Shift the Activity/Impact Gearbox ↓

The new counters are scoped better than the average dashboard.

GitHub’s September 11 changelog adds generally available Copilot usage fields for the dedicated VS Code Agents window. Aggregate reports can include daily active window users plus session and user-message totals; user-level reports can include whether a person used the window and their session/message totals. The fields cover 1-day and 28-day reports.[1]

The most important line is the boundary: GitHub says these fields cover the dedicated Agents window only. They remain separate from editor-window Agent Mode and generic rollups. Missing values can also be absent or null. That is good instrumentation hygiene. A named surface and explicit missingness beat a mystery “AI usage” number.

A counter becomes dangerous when its label is wider than its evidence.

The traffic is about to get heavier.

VS Code 1.137 makes the Agents window a more persistent operating surface. Its release notes describe scheduled automations, quick chats that can become workspace sessions without losing their current request or history, agent messages queued behind busy chats, and an agent host that can expose the same session across multiple VS Code windows.[2]

Those are product capabilities, not proof of benefit or harm. But they change the shape of measurement. A person can trigger recurring work, maintain longer sessions, and coordinate parallel chats. Raw sessions and messages may rise because the surface can now sustain more traffic—not because the work became more useful.

Even a resolved comment is not automatically an outcome.

Another September 11 GitHub update says Copilot code review can automatically resolve its own comments when a later commit addresses the feedback. GitHub also says the reviewer now has broader shell tooling behind its agent firewall, and Lite reviews use an ensemble of agents.[3]

GitHub reports experiment results for that review change: more addressed comments at high, medium, and low severity, plus lower review cost. Those are vendor-reported product measurements, not a general guarantee for every repository. More importantly, “comment addressed” is closer to work than “message sent,” but it still is not the same as fewer escaped defects, less rollback pain, or easier maintenance six months later.

Build a four-gear measurement stack.

  1. Attendance. Who used the named surface? Keep the denominator, reporting window, role, and missing-data rule attached.
  2. Traffic. How many sessions, messages, tool calls, review comments, or automation runs moved through it? Treat volume as workload shape, not value.
  3. Flow. What accepted change, review-cycle reduction, lead-time shift, or completed task moved through the delivery system? Compare like work with a baseline.
  4. Outcome. What happened to escaped defects, rework, rollback frequency, incident recovery, maintainability, and developer understanding?

No single metric gets custody of the story. Adoption without flow may mean curiosity, friction, or training. Flow without outcomes may be faster production of repair work. Outcomes without a baseline may be weather. The measurement contract needs all the labels the dashboard wants to hide.

Refuse surveillance cosplay.

The presence of user-level fields does not make a leaderboard intelligent. GitHub documents access controls around these reports; your organization still owns the decision about why it measures individuals, who can see the data, and what decisions the data may influence.[1]

Default to aggregate team learning. If a metric can affect performance evaluation, compensation, or staffing, document that purpose before collection, involve the people being measured, and prohibit proxy ranking from activity volume. A developer who sends fewer messages because they frame a task well should not look “less engaged” than someone wrestling the agent for forty turns.

Interactive makeover / measurement contract

Activity/Impact Gearbox.

Traditional purpose replaced: one adoption chart with a green arrow. Better version: shift among four evidence planes, lock the surface and baseline, then copy a measurement contract. The tactile carriage, live text, native radios, keyboard path, and reduced-motion mode all report the same state.

Choose the claim your data can carry

The gearbox does not calculate productivity. It exposes which plane you are actually measuring and what must still be attached.

Measurement plane
Claim stop plate / this gear alone cannot prove
AttendanceProductivity, quality, or shipped value.
TrafficUseful work, good prompting, or faster delivery.
FlowFewer defects, less rework, or easier upkeep.
OutcomeCausation without a baseline and confounder check.

Claim load

ATTENDANCE: you can describe adoption in a named surface. You cannot claim productivity.

1/4attendance
Open the three-source instrument panel
[1] GitHub Changelog — “Add VS Code Agents to Copilot usage metrics,” September 11, 2026: dedicated-window fields, 1-day/28-day reports, access, surface boundaries, and optional/null behavior. [2] Visual Studio Code 1.137 release notes, September 9, 2026: scheduled automations, workspace continuation, queued agent messages, and agent-host session behavior. [3] GitHub Changelog — “Auto-resolution and analysis updates in Copilot code review,” September 11, 2026: automatic thread resolution, expanded shell analysis, Lite agent ensemble, and vendor-reported experiment results. [4] Hacker News discussion observed September 12 — discovery context around making and craft only; not evidence for GitHub or VS Code product claims.

Source boundary: the 1–4 dial identifies a selected measurement plane. It is a teaching proxy, not telemetry, a productivity score, a causal estimate, or an employee rating. The interlocks record documentation readiness; they do not validate the underlying data.