Skip to main content

AI Usage Dataset Overview

Written by Tom Williams

The AI Usage dataset holds daily measurements of AI coding assistant usage: who is using AI tools, how much, on what, and to what effect. It brings every AI coding assistant onto a single schema so they can be reported side by side.

Only GitHub Copilot is connected today. The Anthropic Claude, Cursor, OpenAI Platform and OpenAI Codex integrations are coming soon. Fields those providers populate exist in the schema but are empty for now, and each field article says so.

Before building your first widget, read these two articles. AI usage records are shaped differently from every other Keypup dataset, and a report that ignores this will silently double count.

In short: every record covers one UTC day, and two fields say what kind of measurement it is. aggregation_level says whether the record counts a user, a repository or the whole organization. breakdown says which cut of the provider's report it came from. Always filter on one value of each - the cuts overlap, and summing across them counts the same activity several times.

Click on any field to access its documentation.

Identity & period

Field label

Field ID in dataset

_system_id

id

created_at

updated_at

Measurement type

Field label

Field ID in dataset

aggregation_level

breakdown

Actor

Field label

Field ID in dataset

actor_username

actor_type

Surface & language

Field label

Field ID in dataset

surface_ref

surface_name

language

Model

Field label

Field ID in dataset

model_ref

model_name

model_vendor

model_speed

model_context_tier

model_is_auto_routed

Client

Field label

Field ID in dataset

client_name

client_id

client_version

client_plugin

client_plugin_version

Repository

Field label

Field ID in dataset

repository_full_name

repository_name

repository_owner

repository_visibility

Measures - interaction

Field label

Field ID in dataset

interactions

requests

sessions

Measures - tokens

Field label

Field ID in dataset

input_tokens_uncached

input_tokens_cache_read

input_tokens_cache_write

output_tokens

Measures - code outcome

Field label

Field ID in dataset

suggestions_offered

suggestions_accepted

suggestions_rejected

lines_suggested_added

lines_suggested_deleted

lines_added

lines_deleted

commits

pull_requests_created

pull_requests_reviewed

pull_requests_merged

review_suggestions_offered

review_suggestions_accepted

Measures - consumption

Field label

Field ID in dataset

web_search_requests

credits_used

Did this answer your question?