Developers can now replay prefix-cache behavior across 12,002 real agent sessions without accessing the underlying conversations. Harvard MaDSys’ FreeInference Agentic Trace release captures 209 billion input tokens, request timing and tool activity, with block-level identifiers designed for serving-infrastructure experiments.
The dataset covers 16 weeks, from May 17 through September 5, 2026, and represents 267 accounts. The release documentation reports 14 agent harnesses, 1,186,582 model requests and 1,213,347 tool calls from coding and assistant agents using FreeInference.
Infrastructure developers can use these recorded workloads to test KV-cache policies, batching, scheduling and session-aware systems. The dataset records serving behavior; it does not provide readable tasks or training trajectories.
The researchers report three findings: cached input dominates session costs under their pricing assumptions, tool execution offers the largest simulated latency improvement, and a small share of context-changing requests causes most fresh prefill. Each comes with conditions to consider before applying it to other AI products.
What Developers Can Download Now
The release offers two representations through its Hugging Face dataset:
- Raw JSONL traces: About 108 GB decompressed, preserving nested session structure, ordered model requests, tool calls and subagent sessions.
- Flattened Parquet tables: About 9 GB, organized into session, request and tool-call tables. The repository says these are sufficient to reproduce the paper’s analysis and figures.
For initial exploration, the Parquet tables are the more manageable download. Developers can query them with DuckDB or load the dataset through Hugging Face Datasets. Experiments that need the original nested relationships can use the raw traces, which include subagent sessions attached to the tool calls that spawned them.
From a local checkout of the repository, the documented setup and table-download commands are:
make setup
.venv/bin/python -m release download --tables
To download the raw traces instead:
.venv/bin/python -m release download --traces
The repository includes conversion tools, analysis code and prefix-cache simulation artifacts. Its documented reproduction commands include make figures for the paper’s tables and figures, and python -m simulation --week 2026-08-30 to recompute a week of cache simulation.
The dataset and accompanying artifact use CC BY 4.0, allowing sharing and adaptation, including commercial use, with appropriate attribution. The repository asks researchers to cite the accompanying paper, From Requests to Sessions: A Large-Scale Characterization of Human-Driven Agentic Workloads.
The 209-billion-token figure describes aggregate input volume across requests, including context repeatedly sent through an agent session. It is not 209 billion distinct tokens of released text, nor a measure of uncached processing alone.
Replayable Prefix Blocks Preserve Reuse, Not Meaning
The release represents request inputs as chained 16-token prefix blocks, tokenized with tiktoken using o200k_base. The released block identifiers let researchers track shared prefixes without publishing the text behind them.
Researchers can use these blocks to experiment with the key-value cache, or KV cache, which retains attention keys and values from previously processed context. They can replay recorded block accesses, vary capacity or eviction rules, and measure how much prefix reuse a policy preserves.
Session structure adds information that an isolated request log would miss. The traces connect model requests to tool calls and later requests, exposing the periods when an agent pauses inference while another operation runs. Those intervals can inform experiments with cache retention, offloading and prefetching.
The data hub also identifies recorded request arrivals and serving latencies as inputs for batching and scheduling research, and lists LLMLCS export for experiments with libCacheSim.
“Replay” here means replaying workload structure and cache access patterns. The original agent task cannot be rerun from the release because prompts, generated answers and tool-result contents are absent. The standardized token blocks also should not be treated as proof that every anonymized provider uses the same tokenizer or physical cache layout.
Cached Input Dominates Cost Under a Stated Price Ratio
Cached input is the largest cost component in 72% of sessions, the researchers report. Their cost breakdown uses a relative token-price ratio of 1:10:50 for cached input, uncached input and output.
The result depends on that assumption. It does not establish that cached input dominates spending under every model’s pricing, or that uncached-input and output prices are irrelevant.
Repeated context can outweigh a substantial per-token discount. An agent may carry a large conversation history through many tool-use steps. Even if most of that history receives a cached-input rate, repeatedly reading it can contribute more to the modeled bill than the smaller quantities of fresh input or generated output.
For developers comparing model APIs, the practical response is to price a representative session rather than a single request. A cost model should distinguish cached input, fresh input and output, then apply the provider’s actual rates.
Billing and infrastructure also need separate consideration. Lower cache-read prices affect customer cost; better cache placement and retention affect serving work. Improving one does not automatically guarantee an equivalent improvement in the other.
Tool Delays Offer the Largest Simulated Speedup
In the researchers’ latency simulation, halving tool time improves session speed by 1.38×, compared with 1.09× for halving prefill time and 1.16× for halving decode time.
Prefill processes the input context; decode generates output tokens. Optimizing model execution alone, the comparison suggests, can leave much of an agent’s active-session delay untouched.
The simulation conditions exclude user idle time and cap each tool gap at 15 minutes. These speedups were modeled, not measured by deploying faster tools across the service. The trace schema adds another limitation: tool latency is represented through the gap until the next model request. That provides useful operational evidence without fully instrumenting everything happening inside a shell command, browser action or remote service.
Before investing exclusively in faster decoding, infrastructure teams can examine where sessions wait between model requests and which tool categories contribute those waits. They then need to distinguish work that can genuinely run faster from orchestration delay or operations that depend on external systems.
The aggregate result does not imply that every harness is tool-bound. The downloadable traces allow teams to test that question by workload rather than assume one optimization priority for all agents.
Just 4% of Requests Cause Most Fresh Prefill
The researchers report that 4% of requests mutate prior context, yet account for 62.5% of fresh prefill. Append-only requests make up the other 96% of requests and contribute 37.5% of fresh prefill.
The context-mutation analysis assumes infinite KV-cache retention, isolating reuse lost through context changes from reuse lost through eviction.
An append-only request preserves earlier context and extends it. A mutation changes something already present, potentially breaking a reusable prefix. Even an infrequent change can force processing of a much larger section of context.
Sources
- FreeInference Agentic Tracegithub.com
- Hugging Face datasethuggingface.co
- data hubdata.agentic-system.org





