observed workload · replay · evidence

Your workload. Any stack. Replay the difference.

StackReplay reads the usage your coding agents already recorded, then replays that exact demand against another target's real mechanics: rolling windows, allowances, model rules and list prices. What it can establish, it reports. What it cannot, it says so.

Select supported history files or a folder in your browser to load a real workload.

Local-first. Your workload stays in your browser; a share link carries aggregates only.

The Replay instrument

one workload · three targets · engine output, not a mock-up
Execution target
observedObserved workload · recorded before the replay, not generated by it
identityModel identity · resolving observed identifiers against the catalog
targetTarget execution stack · applying the selected target's rules
pressure3 crossingsChronology · advancing the workload through the target's windows
settledResult settled

Result settled

Observed workload: 10,723 synthetic events over 31 days, generated deterministically and replayed by the production engine. Every target, price and limit below comes from the catalog's synthetic example- namespace; nothing here is a claim about a real provider.

simulated outcome76.3%of modelled requests would have fit
simulated target cost$50.00established for this workload

8,180 of 10,723 modeled requests fit (76.3%). 2,543 blocked across 3 crossing windows.

2,543blocked0billed above allowance0undecided1 dimensions partialreplayabilitybounded
01Observed workload
imported workload

This replay covers the imported coding workload only: the events supplied to it, not the whole provider account. Usage outside this workload is not part of the result.

Window31 days covered21 Aug to 20 Sept
Events335 sessions10,723
Models observedevery identifier resolved against the catalog3
Known tokensevery event reports every canonical category214.2M
Uncached input16.3M
Cache read185.6M
Cache write3.9M
Output7.3M
Reasoning1.1M
02Model identity
exact replay · bounded
  • example-large2,370 events · 22.1% of demandas recorded
  • example-medium5,677 events · 52.9% of demandas recorded
  • example-small2,676 events · 25.0% of demandas recorded

No cross-model substitution was applied. Whether every named model is served, and how much demand stayed undecided, is reported under outcomes and evidence.

03Target execution stack
subscriptionsynthetic catalog namespace
ProviderExample Cloud
Planexample-cloud-pro@2026-08-01Example Cloud Pro
Pricefixed month price$50.00
Paid overagethe target does not state whether overage above the allowance is billableunknown
Plan versionwhat this result is pinned toexample-cloud-pro@2026-08-01
Reset phaseNot established: the target's allowance windows mix rolling and calendar behaviour, so no single reset phase is established, and reset-phase sensitivity is not analysed in this milestone
Rules as ofcatalog sha256 73684c134d2… · engine 0.0.4 · methodology 1.4.02026-09-20
04Constraint trace
3 crossings recorded
  • Monthly token allowancecalendar month (UTC) · reject request101.6M / 200.0Mwithin limits
    133.8M tokens attempted0 events rejected
  • 5-hour request windowrolling window of PT5H anchored at first use · latch until reset8,180 / 600exceeded
    10,723 requests attempted2,543 events blocked while latched
  • Large model credit poolcalendar month (UTC) · allow overage$11.62 / $100.00within limits
    $14.81 credit attempted0 events refused (this rule bills the excess instead)
Recorded crossings
5-hour request windowrolling windowfrom 06 Sept, 02:02Z to 06 Sept, 07:02Z1,630 vs 600 included
Window
2026-09-06T02:02:06.709Z to 2026-09-06T07:02:06.709Z
Attempted demand
1,630
Window limit
600
Accepted
600
Affected events
1,030
Above included capacity
not quantified: the rule records the window without splitting the excess
Behaviour
latched until the window reset
Disposition
further requests blocked until the window resets
5-hour request windowrolling windowfrom 06 Sept, 08:02Z to 06 Sept, 13:02Z1,345 vs 600 included
Window
2026-09-06T08:02:44.792Z to 2026-09-06T13:02:44.792Z
Attempted demand
1,345
Window limit
600
Accepted
600
Affected events
745
Above included capacity
not quantified: the rule records the window without splitting the excess
Behaviour
latched until the window reset
Disposition
further requests blocked until the window resets
5-hour request windowrolling windowfrom 06 Sept, 14:02Z to 06 Sept, 19:02Z1,368 vs 600 included
Window
2026-09-06T14:02:30.199Z to 2026-09-06T19:02:30.199Z
Attempted demand
1,368
Window limit
600
Accepted
600
Affected events
768
Above included capacity
not quantified: the rule records the window without splitting the excess
Behaviour
latched until the window reset
Disposition
further requests blocked until the window resets
05Replay outcomes
10,723 events replayed
  • Includedwithin the target's allowance8,180
  • Overageserved, billed above allowance0
  • Blockedrejected or deferred by the rules2,543
  • Unavailableeffective model not served by the target0
  • Unknownevidence insufficient to decide0
06Replay evidence
1 dimension partial
  • Model identityHow much of the workload's model identity is established10,723 of 10,723 events · 214,173,228 of 214,173,228 tokensestablished
  • Usage categoriesHow much of the workload establishes every canonical token category10,723 of 10,723 events · 214,173,228 of 214,173,228 tokensestablished
  • Pricing recordsHow much of the demanded price could be converted to money10,723 of 10,723 eventsestablished
  • Target rulesHow much of the workload the target's rule set describes at all10,723 of 10,723 events · 214,173,228 of 214,173,228 tokensestablished
  • Temporal coverageHow much of the workload lies inside the pinned rule set10,723 of 10,723 events · 214,173,228 of 214,173,228 tokensestablished
  • Translation methodThe token transform applied, if anyNone appliednot applicable
  • Reset phaseWhether the account's allowance reset phase is establishedNot establishedpartial

    the target's allowance windows mix rolling and calendar behaviour, so no single reset phase is established, and reset-phase sensitivity is not analysed in this milestone

+ assumptions this result rests on
  • Admission is atomic across constraints: an event rejected by one rule consumes nothing from any other pool, while attempted demand is still reported per constraint.
  • Plan rules, pricing references and promotions are the snapshot in effect at rulesAsOf; workload chronology inside the simulation uses the historical event timestamps.
  • Consumption uses disjoint canonical token buckets derived from each event's accounting declaration, so overlapping categories are never double counted.
  • A latch_until_reset rule blocks every applicable request until the window that triggered the latch resets.
  • Calendar month windows use each constraint's declared timezone; billing anchors are not supported.
  • The base plan cost is the plan's fixed price; no proration is applied for replay windows shorter than a billing period.
  • Overage is computed per window: units above included capacity in each window are billed at that rule's declared rate.
  • The replay simulates how the target would treat the recorded demand stream: requests after a hypothetical rejection or substitution remain part of the replayed demand, and nothing here models how a person or an agent would have changed behaviour.
  • The target's own windows mix rolling and calendar behaviour, so this replay does not assume one homogeneous reset phase; window outcomes are modelled rather than exact.
  • Window boundaries are half-open: [start, end).
  • Each constraint's windows follow the events it applies to; only served events advance accepted consumption, while rejected events still count as attempted demand.
07Cost counterfactual
established for this workload
Plan pricefixed price of the target plan$50.00
Billed above allowanceno demand was billed above the allowance$0.00
Simulated target costplan price plus billed overage$50.00
  • latch triggered One or more events were blocked by a latch_until_reset rule until its window resets. (2,543 events)

One workload, many targets

one observed demand stream · three targets

Each row is the engine's own result for the same month. Loading a target changes the instrument above; it never changes the workload.

  • Example Cloud ProSubscription · monthly token allowance, request window, one billed credit pool8,180 of 10,723 modeled requests fit (76.3%). 2,543 blocked across 3 crossing windows.example-cloud-pro@2026-08-01 · bounded$50.00simulated target cost
  • Example Cloud StarterTranslatedSubscription · rolling credit pools, with one model substituted under a scenario10,723 of 10,723 modeled requests fit (100%).example-cloud-starter@2026-09-15 · deterministic$20.00simulated target cost
  • Example Cloud Direct APIDirect API · published list prices, per canonical model, no plan10,723 of 10,723 modeled requests fit (100%).example-cloud · deterministic$48.7631 days at list price

no recommendationThese rows describe what each target would have done with this demand. StackReplay does not rank them or tell you which one to buy.

01Workload anatomy
observed · imported workload scope
Time window31 days21 Aug to 20 Sept
Events10,723
Sessionsdeduplicated across harnesses335
SourcesClaude Code, Codex, OpenCode, Command Code, Hermes5
Orchestratorsattribution only: these ran the calls, they did not record tokens1
Known tokensevery event reports every canonical category214.2M
Uncached input16.3M
Cache read185.6M
Cache write3.9M
Output7.3M
Reasoning1.1M
model mixfully resolved
example-large2,370 events47.0M tokens · 21.9% of known tokens
example-medium5,677 events113.8M tokens · 53.1% of known tokens
example-small2,676 events53.3M tokens · 24.9% of known tokens
02Exact versus translated replay
the replay mode states one thing: whether a substitution was applied
Exact replayexact

Example Cloud Pro · example-cloud-pro@2026-08-01

Served within allowance8,180
Blocked by the target's own ruleblocked until the window reset2,543
Not served by the targetevery model is served here0
Crossings recorded3

No cross-model substitution was applied. Whether every named model is served, and how much demand stayed undecided, is reported under outcomes and evidence.

This exact replay still refused 2,543 events: exact says nothing about whether the target serves everything.

Translated replaytranslated

Example Cloud Starter · example-cloud-starter@2026-09-15

Served within allowance10,723
Blocked by the target's own rulenothing was refused or blocked0
Not served by the targetevery model is served here0
Crossings recorded0

Demand recorded against one model was replayed against a substitute under an explicit scenario assumption. That is a counterfactual, not a measurement, and it says nothing about equal capability or equal token consumption.

what the substitution assumed
  • example-largeexample-medium2,370 events substituted

Transform: token preserving · policy demo-scenario-translation@1.0.0. Demand recorded against one model was replayed against a substitute under an explicit scenario assumption. That is a counterfactual, not a measurement, and it says nothing about equal capability or equal token consumption.

A translated replay is a scenario authored for this demonstration. It is not a claim that these two models behave alike, cost alike, or produce the same number of tokens.

Constraint forensics

what the target's own rules did with the demand

A subscription is replayed against its declared rules. Here the month crosses a rolling request window on its heaviest day: 2,543 events would not have been served, and every crossing is inspectable.

03Constraint trace
3 crossings recorded
  • Monthly token allowancecalendar month (UTC) · reject request101.6M / 200.0Mwithin limits
    133.8M tokens attempted0 events rejected
  • 5-hour request windowrolling window of PT5H anchored at first use · latch until reset8,180 / 600exceeded
    10,723 requests attempted2,543 events blocked while latched
  • Large model credit poolcalendar month (UTC) · allow overage$11.62 / $100.00within limits
    $14.81 credit attempted0 events refused (this rule bills the excess instead)
Recorded crossings
5-hour request windowrolling windowfrom 06 Sept, 02:02Z to 06 Sept, 07:02Z1,630 vs 600 included
Window
2026-09-06T02:02:06.709Z to 2026-09-06T07:02:06.709Z
Attempted demand
1,630
Window limit
600
Accepted
600
Affected events
1,030
Above included capacity
not quantified: the rule records the window without splitting the excess
Behaviour
latched until the window reset
Disposition
further requests blocked until the window resets
5-hour request windowrolling windowfrom 06 Sept, 08:02Z to 06 Sept, 13:02Z1,345 vs 600 included
Window
2026-09-06T08:02:44.792Z to 2026-09-06T13:02:44.792Z
Attempted demand
1,345
Window limit
600
Accepted
600
Affected events
745
Above included capacity
not quantified: the rule records the window without splitting the excess
Behaviour
latched until the window reset
Disposition
further requests blocked until the window resets
5-hour request windowrolling windowfrom 06 Sept, 14:02Z to 06 Sept, 19:02Z1,368 vs 600 included
Window
2026-09-06T14:02:30.199Z to 2026-09-06T19:02:30.199Z
Attempted demand
1,368
Window limit
600
Accepted
600
Affected events
768
Above included capacity
not quantified: the rule records the window without splitting the excess
Behaviour
latched until the window reset
Disposition
further requests blocked until the window resets

A Direct API target has no allowance to exceed, so it is never given one: selecting it in the instrument replaces this trace with pricing applicability and the provider's own list price, which is everything a metered target can establish about a workload.

Replay evidence

independent dimensions · no combined score

Each dimension answers its own question about the same replay. They are reported separately on purpose: a single confidence figure would hide which part of the answer rests on what.

04Replay evidence
1 dimension partial
  • Model identityHow much of the workload's model identity is established10,723 of 10,723 events · 214,173,228 of 214,173,228 tokensestablished
  • Usage categoriesHow much of the workload establishes every canonical token category10,723 of 10,723 events · 214,173,228 of 214,173,228 tokensestablished
  • Pricing recordsHow much of the demanded price could be converted to money10,723 of 10,723 eventsestablished
  • Target rulesHow much of the workload the target's rule set describes at all10,723 of 10,723 events · 214,173,228 of 214,173,228 tokensestablished
  • Temporal coverageHow much of the workload lies inside the pinned rule set1,473 of 10,723 events · 28,946,093 of 214,173,228 tokenspartial

    9250 event(s) fall outside the pinned plan version's effective window (2026-09-15 to open)

  • Translation methodThe token transform applied, if anytoken-preserving · demo-scenario-translationestablished
  • Reset phaseWhether the account's allowance reset phase is establishedEstablishedestablished
+ assumptions this result rests on
  • Admission is atomic across constraints: an event rejected by one rule consumes nothing from any other pool, while attempted demand is still reported per constraint.
  • Plan rules, pricing references and promotions are the snapshot in effect at rulesAsOf; workload chronology inside the simulation uses the historical event timestamps.
  • Consumption uses disjoint canonical token buckets derived from each event's accounting declaration, so overlapping categories are never double counted.
  • A latch_until_reset rule blocks every applicable request until the window that triggered the latch resets.
  • Cross-model translation is an explicit scenario assumption: the substitute model is not an alias of the recorded model, and nothing here claims equal capability, quality, output length or tool behaviour.
  • The base plan cost is the plan's fixed price; no proration is applied for replay windows shorter than a billing period.
  • The replay simulates how the target would treat the recorded demand stream: requests after a hypothetical rejection or substitution remain part of the replayed demand, and nothing here models how a person or an agent would have changed behaviour.
  • The translation preserves the recorded token quantities: the substitute model is assumed to consume the same input, cache, output and reasoning amounts. No empirical conversion ratio or equivalence is applied or implied.
  • Window boundaries are half-open: [start, end).
  • Each constraint's windows follow the events it applies to; only served events advance accepted consumption, while rejected events still count as attempted demand.
The same targets over a workload that keeps two things unknownpartial evidence

488 events in this companion sample report an incomplete token category set, and a few name a model no catalog entry maps (gpt-6-astra-omen). Consumption constraints become unknown rather than silently satisfied, and the unresolved identifier keeps its own lane.

Request coverage is unknown: the replay could not decide every event. 109 left undecided by the evidence.

06Replay evidence
5 dimensions partial
  • Model identityHow much of the workload's model identity is established449 of 488 eventspartial

    39 event(s) use identifiers no catalog source establishes; 74 event(s) report no complete token accounting, so the usage-weighted fraction is limited to the remaining events

  • Usage categoriesHow much of the workload establishes every canonical token category414 of 488 eventspartial

    74 event(s) do not report every canonical token category, so no non-overlapping token denominator exists for the whole workload

  • Pricing recordsHow much of the demanded price could be converted to money379 of 449 eventspartial

    part of the priced demand could not be converted to money

  • Target rulesHow much of the workload the target's rule set describes at all449 of 488 eventspartial

    39 event(s) use models or identifiers the target's rules do not mention

  • Temporal coverageHow much of the workload lies inside the pinned rule set70 of 488 eventspartial

    418 event(s) fall outside the pinned plan version's effective window (2026-09-15 to open)

  • Translation methodThe token transform applied, if anytoken-preserving · demo-scenario-translationestablished
  • Reset phaseWhether the account's allowance reset phase is establishedEstablishedestablished
+ assumptions this result rests on
  • Admission is atomic across constraints: an event rejected by one rule consumes nothing from any other pool, while attempted demand is still reported per constraint.
  • Plan rules, pricing references and promotions are the snapshot in effect at rulesAsOf; workload chronology inside the simulation uses the historical event timestamps.
  • Consumption uses disjoint canonical token buckets derived from each event's accounting declaration, so overlapping categories are never double counted.
  • A latch_until_reset rule blocks every applicable request until the window that triggered the latch resets.
  • Cross-model translation is an explicit scenario assumption: the substitute model is not an alias of the recorded model, and nothing here claims equal capability, quality, output length or tool behaviour.
  • The base plan cost is the plan's fixed price; no proration is applied for replay windows shorter than a billing period.
  • The replay simulates how the target would treat the recorded demand stream: requests after a hypothetical rejection or substitution remain part of the replayed demand, and nothing here models how a person or an agent would have changed behaviour.
  • The translation preserves the recorded token quantities: the substitute model is assumed to consume the same input, cache, output and reasoning amounts. No empirical conversion ratio or equivalence is applied or implied.
  • Window boundaries are half-open: [start, end).
  • Each constraint's windows follow the events it applies to; only served events advance accepted consumption, while rejected events still count as attempted demand.

Cost counterfactual

the category model the provider actually bills

A metered target is priced category by category at the provider's published rates: uncached input, cache read, cache write, output and reasoning, each on its own disjoint token quantity. A category with no documented rate is not priced and never treated as free, so a total only appears when every served event could be priced.

05Cost counterfactual
established for this workload
Uncached input16.3M
Cache read185.6M
Cache write3.9M
Output7.3M
Reasoning1.1M
Counterfactual at list price10,723 of 10,723 served events priced$48.76

The same month against the example subscription targets costs $50 and $20.00 respectively. Those are plan prices plus any billed overage, which is a different basis from a metered list price; the two are shown side by side, never merged into one comparison.

What a replay is, and what it is not

  • It is a simulation of documented target mechanics over your recorded demand stream, with the assumptions it used printed beside the result.
  • It reports what it could determine, what it could not, and how much of the answer is established, dimension by dimension.
  • It runs in your browser. No workload content is uploaded, and there is no account.
  • It is not a bill, and no figure here is a provider's statement of account.
  • It does not recommend a target. Comparison here is descriptive; ranking is not built.
  • It does not model how you or your agents would have changed behaviour after a rejection, so it never claims a workload would have been run the same way elsewhere.

What leaves your machine

by default: nothing

StackReplay reads the files you select and processes them locally in a browser worker. Raw workload files are not uploaded. You can replay a temporary import, or choose to save its normalized workload in this browser for later.

if you share a result

A share link carries aggregates only: counts, token totals, the target, coverage, constraint summaries and versions. Never events, session or project identifiers, repository names, paths, prompts, responses or file names.

Catalogued plans

All plans

Where the project is

The replay engine, the adapters and the local browser application are built and tested. The hosted services that the specification describes (sync and sharing across devices) are not built and are not presented here as if they were. This site publishes what exists.