Tuesday, September 15, 2026Verified technology journalism

The moving target, explained: why the AI you pay for is never the same product twice

AI subscriptions sell an artifact and deliver a service: the vendor keeps control of which model answers, how much of it you get, and how it is served, so quality can change after you subscribe without anything being announced. The four failure modes: silent model retirement, allowance erosion at a constant price, identical weights scoring differently by provider, and retention rules changing under your data, plus what to watch instead of release notes.

The moving target, explained: why the AI you pay for is never the same product twice

An AI subscription sells you an artifact and delivers a service. The artifact is the thing you believe you bought: a named model, at a named price, that behaved a specific way the week you evaluated it. The service is what actually arrives: whichever model the vendor routes to, at whatever allowance current policy allows, on whatever hardware the provider selected, under whatever retention terms now attach to what you typed. The gap between them is where the product changes after you pay.

That gap has a measurable shape. Claude flagship models released in 2024 lived 16 to 22 months from launch to API retirement; every flagship Anthropic released in 2025 lived 12 to 13 months 1. Four levers explain it, and the vendor holds all four.

One purchase, four moving parts

The mechanism fits in one sentence: the vendor keeps four controls that change the product without touching the price tag. Which model answers is the retirement lever. How much of it you get is the allowance lever. How it is served is the serving lever. What happens to your data afterward is the terms lever. Pull any one and the experience changes while the receipt stays identical.

None of it is secret. Every change in this piece has a paper trail in a deprecation table, a help-center notice, or a routing document. The trail simply lives on surfaces a chatbot subscriber has no routine reason to open. One calendar year of the pattern, all four levers firing:

  • May 17, 2026: Google switches Gemini from daily prompt limits to compute-based usage, five-hour refreshes under a weekly cap; Google's own limits page carries the dated notice 2
  • May 19, 2026: Google splits AI Ultra, adding a $100 entry tier while the previous $250 plan drops to $200 with what Google calls "the exact same capabilities" 2
  • June 2, 2026: Anthropic's help center, date-stamped that day, documents Claude's five-hour session limits with weekly limits on top 2
  • June 15, 2026: Claude Sonnet 4 and Claude Opus 4 reach their API retirement dates, 13 months after launch 3
  • June 18, 2026: Gemini CLI stops serving Google AI Pro and Ultra subscribers, per the maintainers' May 19 announcement 2
  • September 1, 2026: Claude Fable 5.1 reaches general availability in GitHub Copilot with prompt retention attached as a condition of use 4
  • December 11, 2026: OpenAI's scheduled removal date for its GPT-5 and o3 snapshots, deprecated with notice on June 11 5
  • December 31, 2026: expiry of the exemption letting approved enterprises run Fable 5.1 in Copilot without retention 4

The model you tested on stops existing

Retirement is the loudest lever and the only one with a published schedule. Requests to a retired model fail outright 3. What the schedules show is compression: Claude flagship lifespans ran 22 months (Claude 3 Opus), then 16 (Claude 3.5 Sonnet), then a flat 12 (Claude 3.7 Sonnet and Claude Opus 4.1, each retired exactly twelve months after launch); nine Claude versions were retired between October 2025 and August 2026 1. Google's preview-tier models are scheduled for a median of about six months, roughly 40 percent of a Google GA model's scheduled lifespan; Google's own table flags its dates as the earliest a model might go, so those spans are floors 1. The notice floors differ by lab:

  • OpenAI: at least 6 months notice for generally available models and at least 3 for specialized variants; preview models "may be retired with much shorter notice, such as 2 weeks" 5
  • Anthropic: at least 60 days for publicly released models, and its three 2026 notices ran 60, 61 and 62 days 1
  • Google: no minimum notice period stated on its deprecations page 1

Two wrinkles make retirement quieter than the tables suggest. OpenAI notifies customers actively using the model by email and documents the change on the deprecations page, "along with blog posts for larger changes" 5; a consumer subscriber is not that audience. And Anthropic's dates bind only Anthropic-operated platforms: partner platforms like Amazon Bedrock and Google Cloud "set their own retirement schedules," so the same model dies on different days depending on where you run it 3. The subscription side of the same lever is deployment dependence: GPT-6 Astra reached enterprise customers first while Plus and Pro subscribers waited, and OpenAI's make-good was one banked reset per day without access, a currency OpenAI itself mints and values 2.

The price holds while the allowance moves

Allowance erosion never presents as a price change, because it is not one. ChatGPT Plus has read $20 since February 2023, Google's paid tier has listed $19.99 since February 2024, and Claude Pro lists at $20 2. The May 17 switch to compute-based usage at Google, the June 2 dual clock at Anthropic, the June 18 loss of Gemini CLI access for paying AI Pro and Ultra subscribers: all landed at prices that did not move 2. The two vendors converged within weeks on the same architecture, a five-hour window, a weekly ceiling, and credits to pay past the cap, which Anthropic bills at standard API rates 2. Anthropic reserves the move in its own words: "Price and plans are subject to change at Anthropic's discretion" 2. At the margin, the all-inclusive subscription becomes a base fee plus a meter: the mobile-carrier model.

Identical weights, a different machine

The subtlest lever requires no model change at all. OpenRouter, a service that routes one API request across many model providers, documents the mechanics plainly. By default it load-balances across the top providers "to maximize uptime," weighting selection by the inverse square of price; in the service's own worked example, a provider charging a third as much is nine times as likely to answer first 6. The same open-weight model name is served by different providers at different quantization levels, which is why the router exposes a quantization filter (int4, int8) for requests 6. Fallbacks are on by default, so a provider with outages in the last 30 seconds loses your traffic to another one, request by request 6.

The implication for buyers: the weights are the artifact, the serving stack is the service. Buy access to a model through a routing layer and the deployment that answers is a runtime decision made on price and uptime, not a property of what you purchased. From the vendor's side, nothing changed: the model string you asked for is identical. Your outputs still move.

The terms move under your data

Retention is the lever aimed at your inputs rather than your outputs. ProvenBrief's retention tracker, release one, maps five verified model-and-tier cells across Anthropic's commercial surfaces; four of the five retain prompts by default 4. Claude Enterprise conversations are retained indefinitely unless an owner sets a custom period, floored at 30 days, a control that shipped March 16, 2026 4. Claude Fable 5.1 in GitHub Copilot cannot run at all unless Anthropic retains prompts and outputs to operate its safety classifiers, and zero-retention access exists only through a time-bound exemption dated through December 31, 2026 4. Covered Models under a Business Associate Agreement hold a 30-day retention floor with zero-retention unavailable altogether 4. The pattern to internalize: a retention setting is only as durable as the policy underneath it, and the policy has dates.

Some of this is cost structure, not trickery

The honest version carries the vendors' side. Anthropic has put the economics on the record: the cost of keeping models available for inference scales roughly linearly with the number of models served, and its own documentation says retirement happens to ensure capacity for new releases 1 3. Serving a frontier model under a flat, unlimited $20 plan is a business that only works with a meter or a shrinking quota; the dual clocks are partly that cost structure arriving in the product. Pinning has its own price: Anthropic's page says deprecated models are "likely to be less reliable," and legacy models stop receiving updates, and OpenAI's soft landing is paid dedicated capacity OpenAI may offer after a model's shutdown date 3 5. Freeze a version and you freeze yourself out of the fixes. The moving target is partly the physics of the business. The failure is marketing it as a statue.

What to watch instead of release notes

You cannot stop the target moving, but you can watch the surfaces where it actually moves, and none of them is the product blog:

  • Pin dated snapshots where the API allows. OpenAI's dated snapshots run 14 to 20 months of life 1; on OpenRouter you can pin providers with an order or allowlist, filter quantizations, and switch fallbacks off, trading failover for determinism 6
  • Read the deprecation tables, not the roadmap. OpenAI's deprecations page and Anthropic's model deprecations page are the real forward schedule, and Google's deprecations page states no minimum notice at all 5 3
  • Run a fixed eval set on a schedule: the same prompts, the same rubric, every month. It is the only instrument here that catches allowance erosion and serving variance, the two levers with no deprecation row; Anthropic itself advises testing replacement models well before retirement dates 3
  • Diff the allowance terms at renewal. The vendors' own limits pages carry dated change notices, as Google's May 17, 2026 notice shows; the notice, not the price, is where the change lands 2

The working assumption that fits all four levers: you do not subscribe to a model. You subscribe to a slot, and the vendor decides what fills it this week. The model card is marketing. The deprecation table is the contract.

Cite this story

ProvenBrief (2026). "The moving target, explained: why the AI you pay for is never the same product twice." ProvenBrief. https://provenbrief.com/story/the-moving-target-explained-why-the-ai-you-pay-for-is-never-the-same-product-twi

Free to quote and link with attribution. Republishing in full or AI-training use requires a license.

Verified49 factual claims in this story were independently checked against primary sources before publication. Read our editorial standards.

Get the next brief in your inbox

One weekly email. Every claim verified against primary sources before we hit send.

Produced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.