Task-to-Model Routing
Match each AI task to the cheapest model capable of doing it well
- Difficulty
- Advanced
- Time to result
- ~months to results
- Steps
- 6
- Confidence
- 88%
Task-to-Model Routing treats model selection as an organizational design problem rather than an ad hoc user choice. The company inventories recurring tasks, classifies their complexity and risk, and defines the minimum quality each task requires. It then benchmarks candidate models and routes each task to the least expensive option that reliably clears the threshold. Simple transformations may use a cheap model, while ambiguous reasoning or high-risk work receives a stronger one. Stable, high-volume workflows may justify a fine-tuned open-source model adapted to company context. Costs, output quality, and exceptions are monitored so routing rules can change as models improve. This replaces the default behavior of choosing the “best” model for everything with a governed portfolio that spends premium tokens only where additional capability creates meaningful value.
Origin
Extracted from Marketing Against The Grain during a discussion of rising token budgets and the difficulty of asking ordinary employees to choose among models.
Core principles
- 01Model capability should match task complexity
- 02Simple tasks should not consume premium-model budgets
- 03Users should not bear the full routing burden
- 04Specialized models can outperform wasteful universal defaults
- 05Routing rules must preserve required output quality
How to run it
- 1
Inventory recurring tasks
List the repeated AI tasks performed across teams and estimate their frequency and current token cost.
Pro tip Group similar prompts by the business job they perform.
Watch out Do not optimize rare tasks before identifying the major sources of spend.
- 2
Classify complexity and risk
Rate each task by reasoning difficulty, context requirements, consequence of error, and need for specialized knowledge.
Pro tip Separate reversible drafting from consequential decisions.
Watch out A short prompt is not necessarily a simple or low-risk task.
- 3
Define quality thresholds
Specify the minimum acceptable accuracy, format compliance, latency, and review burden for each task class.
Pro tip Use representative test cases and explicit pass criteria.
Watch out Cost comparisons are meaningless when candidate outputs do not meet the same standard.
- 4
Benchmark candidate models
Test cheap, premium, and where appropriate open-source models against the task set. Compare total cost, including retries and human correction.
Pro tip Measure effective cost per accepted output.
Watch out Do not route solely from advertised benchmark scores.
- 5
Create the routing map
Assign each task to the cheapest model that consistently satisfies its threshold, with escalation rules for failures or unusual inputs.
Pro tip Automate routing so ordinary users do not need to understand model economics.
Watch out Never send high-risk exceptions silently to an underpowered model.
- 6
Review and refine
Monitor quality, cost, and failure patterns, then update mappings as workloads and models change. Evaluate fine-tuning when a stable task has enough volume to justify it.
Pro tip Re-benchmark after major model or pricing changes.
Watch out Static routing rules become obsolete quickly in a changing model market.
In the wild
A content organization routes metadata formatting and headline variations to a low-cost model, research synthesis to a stronger reasoning model, and final claims through human review. Each route is selected from benchmarked quality thresholds rather than employee preference.
→ The organization lowers token expenditure while preserving quality where capability matters.
A company has a stable, high-volume task that classifies support tickets into a fixed taxonomy. It benchmarks a fine-tuned open-source model against a premium general model and adopts the specialized option after it meets the same acceptance threshold at lower cost.
→ A predictable recurring task moves to a cheaper specialized model.
Common mistakes
Using the best model for everything
Premium capability is wasted when a cheaper model can meet the task's actual requirements.
Delegating routing to every user
Most employees will choose based on perceived quality or convenience rather than systematic cost-performance evidence.
Optimizing token price alone
Cheap outputs that require retries or extensive correction may have a higher effective cost.
Is it for you?
Best for
It is best for organizations with substantial recurring AI volume across tasks of widely different complexity and risk.
Not ideal for
It is not ideal for small teams whose low usage does not justify the benchmarking, routing, and maintenance overhead.
From the transcript
“At some point you're gonna have to say what model should I use?”
“This is a simplistic task. I should use a really cheap model.”
“I think companies will start to integrate open source models and fine-tune them for their own companies and then you'll actually have tasks mapped to…”
From the episode
They Spent $150,000 on AI Tokens (And Got Nothing)