Perspective · AI Leadership

Every AI workload must earn its model tier.

Prove the work with enough capability. Measure what it earns. Then decide whether it still deserves Frontier, General or Specialised capability. The model should follow the business case, not lead it.

There is a quiet cost appearing inside AI programmes. A team proves a workload on the strongest model available, puts it into production, and moves on. Six months later the model market has changed, but the workload has not. The organisation keeps paying for capability it may no longer need.

The initial decision may have been completely rational. When a team is trying to determine whether AI can perform a piece of work, deliberately constraining the model can distort the test. If a Frontier model is the best available way to prove the workload, use it.

The mistake comes later. A successful proof is often treated as a permanent model decision.

The Techshin Partners view

Prove the workload with capability. Then make the capability earn its place.

Before you read further, ask five questions.

You do not need to know the token price of every model to know whether you have a model selection problem. Start with visibility.

Can you name the AI workloads that are operating today?Not the tools. The actual pieces of work being performed.
Do you know why each workload uses its current model?Was the model deliberately selected, or did it simply remain after the proof?
Do you know what business value each workload creates?Time saved, capacity released, quality, margin, revenue, risk or customer impact.
Has the model decision been reviewed since the workload went live?The workflow may be stable while the market underneath it changes quickly.
Could you explain the capability premium you are paying for?If Frontier costs more, what capability does that additional spend buy and does the workload need it?

If several of those questions are difficult to answer, the issue is not necessarily overspending. The issue is that the organisation cannot yet explain whether the spend is justified.

The hidden problem is Capability Debt.

Technology leaders already understand technical debt. AI introduces another form of debt.

Capability Debt is the continuing cost of using more model capability than a workload now requires.

It happens naturally. A model is chosen when the workload is uncertain. The team proves the process. The prompts improve. The SOP becomes clearer. Inputs become more structured. Guardrails get stronger. The work itself becomes more predictable. At the same time, models elsewhere in the market become cheaper, faster and more capable.

Yet the production workload remains exactly where it started.

Capability Debt

Frontier today can become General tomorrow without the workflow changing at all.

This does not mean every workload should move down a model tier. Some workloads will continue to justify Frontier capability. The point is that the decision should be earned by evidence, not inherited from history.

Three capability tiers. One business question.

Specialised

Narrow capability for repeatable, bounded work. Often attractive where cost, latency, privacy, local deployment or highly specific behaviour matters.

General

Strong everyday reasoning and language capability for understood business workloads with manageable variation and clear operating controls.

Frontier

The strongest broadly available capability for ambiguity, complex synthesis, deeper reasoning, difficult exceptions and workloads where quality dominates price.

These tiers are deliberately not tied to vendor names. Models move. Product names change. Providers merge, retire and reposition their offerings. The useful thing to record is the capability the workload requires.

The question is not, which model do we like?

The question is:

The decision question

What capability are we paying for, and is that capability worth the premium?

EARN. A repeatable way to make the decision.

The EARN framework separates proving the workload from optimising it. That matters because optimisation introduced too early can prevent the organisation from discovering value in the first place.

E

Establish

Prove that the workload works and establish the business value before constraining model capability.

A

Assign

Assign the capability tier the workload currently requires based on quality, variability, consequence and review.

R

Refine

Compare the economics and test whether a lower tier can deliver the same required outcome reliably.

N

Next Review

Set the trigger or date to test the decision again as price, capability and operating conditions change.

E. Establish the value before you optimise the cost.

A workload should first prove that it deserves to exist. That means measuring what it changes in the business, not simply whether the model produced an impressive output.

  • How many times does the workload run?
  • How much human time does it remove or release?
  • What is that time worth?
  • Does it improve quality, response time or consistency?
  • Does it reduce risk or avoid costly errors?
  • Does it protect margin or create revenue?
  • What human oversight does it still require?

If the workload earns $20,000 of value and costs $400 to operate on Frontier, the first question is not whether another model can save $200. The first conclusion is that the workload has earned the right to continue.

A. Assign autonomy and capability separately.

One of the easiest mistakes is to combine two different decisions.

Autonomy answers how independently the workload operates. Capability answers how much model strength the workload requires.

Decision Question Possible outcome
Autonomy How much of the work should AI be allowed to perform without human intervention? Augment, Automate, Agentic
Capability What level of model capability is required to perform that work reliably? Specialised, General, Frontier

An Agentic workload does not automatically require a Frontier model. An Augment workload does not automatically belong on a cheaper model. These are separate decisions and both need evidence.

Capability pressure rises when...

The work contains ambiguity, broad context, difficult exceptions, deep reasoning, significant consequence or long periods between human review.

Capability can often reduce when...

The work is narrow, stable, well documented, structured, highly repeatable, easy to evaluate and protected by strong guardrails.

R. Refine the economics. Tokens come later.

Token economics matters. It simply should not lead the conversation.

The cost of a workload is broader than the price of inference. Work through the economics in this order.

01
Business ValueWhat does the workload earn, protect, improve or release?
02
Human EconomicsHow much human effort is removed, redirected or made more productive?
03
Operational EconomicsIntegration, infrastructure, monitoring, review, governance and support.
04
Token EconomicsInput, output, cached tokens, inference, API or local compute costs at the actual workload volume.
05
Model OptimisationCan another capability tier deliver the same agreed quality and operating outcome for less?
Token economics

Saving $40 in model cost while giving away $4,000 of business value is not optimisation.

Once the value is understood, token economics becomes useful. Teams can model monthly run volume, typical input and output size, cost by model tier, latency, human review effort and the cost of failures or rework.

At that point the discussion becomes commercially useful. A General model may cost one quarter of Frontier but run more slowly. A Specialised model may reduce cost again but require tighter inputs. Frontier may remain the correct choice because the value of additional reasoning or accuracy is worth far more than the premium.

N. The model decision needs an expiry date.

The final part of EARN is deliberately not called Now. It is Next Review.

AI model capability is moving too quickly for a production decision to be treated as permanent. The workload may not change for a year while the available model economics change several times.

Set a review date or a review trigger when the workload is approved.

  • A new model reaches the required quality bar.
  • Pricing changes materially.
  • The provider merges or retires the model.
  • Volume changes enough to alter the economics.
  • Latency, quality or drift moves outside the agreed range.
  • Privacy, security or data location requirements change.

This is not a new use case review. The business problem does not need to be reopened every time. It is a capability review against an already understood workload.

A simple example makes the economics obvious.

Worked example

Customer service response drafting

A team proves the workload using Frontier. It runs 4,000 times each month and releases enough human capacity to create an estimated $18,000 of monthly business value.

$18,000Monthly value
$320Frontier cost
FrontierTier earned today

The workload is commercially attractive even at Frontier. There is no reason to delay deployment while trying to save a small fraction of the value.

Three months later, a General model reaches the same agreed quality bar. The same workload costs $75 per month. The organisation retests it, confirms the result and moves the workload down a tier.

What changed?

The workflow did not change. The market did.

The decision should become part of governance.

EARN is not useful if it lives only in someone's head. The outcome belongs with the use case, the SOP and the governance record.

For each production workload, record:

EARN decision record
Workload
Business value
Autonomy
Model tier
Capability premium
Quality bar
Decision owner
Next review

That record creates a much stronger governance conversation than simply writing down which model is in use. It explains why the workload is operating at that level and when the decision should be challenged again.

The goal is not smaller models. It is better decisions.

Some workloads will remain on Frontier. They should. Others will move to General. Some will eventually belong on highly Specialised models running close to the work.

EARN does not predetermine the answer.

It creates a repeatable way to prove value, separate autonomy from capability, understand the complete economics and challenge yesterday's model decision against today's market.

The EARN principle

Every AI workload should earn the capability you are paying for.

That is how organisations avoid optimising too early and avoid paying Capability Debt forever.

Put the thinking into practice

The AI Framing Sprint helps leadership teams establish the governance, operating artefacts and measurement needed to make AI decisions deliberately rather than by default.

Read more about the AI Framing Sprint →
TECHSHIN PARTNERS
Technology  Leadership
A technology leadership practice working alongside CEOs and their leadership teams. Three practice areas. One standard underneath. Currently helping businesses get on the right track with AI.
Same people. Same budget. Different results.