Skip to main content

AI Usage > Input tokens (cache read)

Written by Tom Williams

Dataset: AI Usage

Entity: AI Usage Measurement

Field ID: input_tokens_cache_read

Type: Number

Description: The number of input tokens served from the prompt cache.

Not available yet. GitHub Copilot does not publish this measurement, so it is empty on every record you can query today. It is populated by Anthropic Claude, Cursor and OpenAI Platform, which are coming soon.

Source: App

Transformation logic: N/A

App Mapping

GitHub Copilot

N/A

Anthropic Claude (coming soon)

tokens.cache_read on Claude Code records; cache_read_input_tokens on direct API traffic

Cursor (coming soon)

the sum of tokenUsage.cacheReadTokens across the record's usage events

OpenAI Platform (coming soon)

input_cached_tokens

OpenAI Codex (coming soon)

N/A

Reporting Use Cases

The Input tokens (cache read) field counts prompt tokens that were served from cache rather than reprocessed. They are substantially cheaper, so the cache ratio is one of the few AI efficiency levers a team controls.

  • Cache efficiency: Cache reads as a share of all input tokens shows how well long-running sessions are reusing context. A falling ratio usually means sessions are being restarted more often.

  • Explaining flat costs under rising usage: A team whose cache ratio is improving can grow usage substantially without a matching rise in consumption.

  • Only three providers report it: Copilot and Codex publish no cache split.

Did this answer your question?