Skip to main content

Definition: AI Usage Record Structure

Written by Tom Williams

The AI Usage dataset holds one record per measurement published by an AI provider for one UTC day. Unlike the other Keypup datasets, a record is not a thing (an issue, a commit, a review) but a count - and the same day's activity is published by every provider in several overlapping cuts at once.

Only GitHub Copilot is connected today. The Anthropic Claude, Cursor, OpenAI Platform and OpenAI Codex integrations are coming soon, and fields only they populate are empty for now. See Definition: AI Provider Coverage & Limitations.

Two fields say what kind of measurement a record is, and you almost always need both in your filters.

Aggregation level - what the record counts over

The aggregation_level field says whose activity the record covers.

  • USER the measurement covers one person or one service account.

  • REPOSITORY the measurement covers one repository.

  • ORGANIZATION the measurement covers the whole account.

These are not three views of the same numbers, and they are not interchangeable. Providers publish genuinely different reports at each level: GitHub Copilot's repository report carries only pull request activity, with no actor, no model and no surface, and it can contain data on days with no IDE activity at all. An organization record is published by the provider itself, not computed by summing its user records, and the two will not match.

Always filter on a single aggregation level. A report that mixes them adds a day's per-user counts to that same day's organization total and roughly doubles every figure.

Breakdown - which cut of the report

A single provider report contains several cuts of the same day's activity. One GitHub Copilot user report contains six: a day total plus five breakdown arrays, each slicing the same activity differently. The breakdown field says which cut a record came from.

Each value names the dimensions the record is split by, so it also tells you which dimensions carry a value and which are empty:

Breakdown

Populated

Empty

TOTAL

the subject's day total

surface_ref, model, client, language

SURFACE

surface_ref, surface_name

model, language

SURFACE_MODEL

surface_ref, model

language

CLIENT

client

surface_ref, model, language

LANGUAGE_SURFACE

language, surface_ref

model

LANGUAGE_MODEL

language, model

surface_ref

Always filter on a single breakdown. Summing across breakdowns counts the same activity several times over - a user's day total and their per-surface records describe the same work.

The period is always one UTC day

Every AI provider publishes usage per calendar day, so created_at is the UTC midnight of the day the measurement covers rather than a record creation timestamp. There is no start and end: the period is always that day.

This has two consequences for reporting. Grouping by anything finer than a day returns nothing useful, and a timezone other than UTC will split a provider's day across two of your days.

Records are restated after the fact

AI usage telemetry arrives late. Clients upload asynchronously, and no provider declares a past day final - a day re-read a week later can legitimately carry larger numbers. Keypup reloads a rolling window of recent days and updates the affected records in place, moving updated_at when a figure actually changes.

Recent days are therefore the least stable part of any AI usage chart. Expect the last few days of a trend to climb as telemetry catches up.

Picking a starting filter

For most reports, start from:

  • Per-developer adoption: aggregation_level = USER and breakdown = TOTAL, or breakdown = SURFACE when you want the split by where the work happened.

  • Token consumption per model: aggregation_level = USER and breakdown = SURFACE_MODEL, which is the only cut carrying model attribution.

  • Repository-level pull request impact: aggregation_level = REPOSITORY, which only GitHub Copilot publishes.

Did this answer your question?