Build for the Model 18 Months Ahead
Prepare AI use cases now so better models unlock them instantly
- Difficulty
- Advanced
- Time to result
- ~months to results
- Steps
- 6
- Confidence
- 97%
Start with a business use case that would become highly valuable if AI reliability improved over the next 12 to 24 months. Instead of postponing the entire project, construct the surrounding infrastructure now: data access, orchestration, interfaces, safeguards, and measurement. Connect the best current model as a minimum viable implementation, even if its output is not yet suitable for customers. This exposes architectural weaknesses and clarifies the quality threshold the model must cross. Keep the use case available for repeated testing, then substitute each major model release into the same system. Because the workflow is already operational, a capability improvement can translate into an immediate product improvement rather than initiating a new development cycle.
Origin
Kieran Flanagan described this approach while discussing how companies should prepare for rapidly improving language models on Marketing Against The Grain.
Core principles
- 01AI models will improve faster than most implementation teams can react.
- 02Current model limitations should not define the long-term product vision.
- 03Infrastructure built early creates an advantage when capabilities improve.
- 04A weak prototype can validate architecture before it validates output quality.
- 05Use cases must provide value beyond merely wrapping a model provider.
How to run it
- 1
Choose a future-worthy use case
Find a valuable workflow that AI could transform but cannot yet execute reliably enough. Prioritize durable business value rather than a novelty that the model provider may absorb.
Pro tip Look across the entire business rather than restricting the search to obvious content-generation tasks.
Watch out A generic wrapper around a model is vulnerable to being displaced by the model provider.
- 2
Define the future capability threshold
Describe what the model must accomplish in 12, 18, or 24 months for the use case to become viable. Make reliability, latency, cost, and safety expectations explicit.
Pro tip Frame the threshold as observable acceptance criteria rather than a vague expectation that AI will improve.
Watch out Do not assume every capability will improve at the same rate.
- 3
Build the surrounding infrastructure
Implement the required data pipelines, workflow logic, interfaces, permissions, and measurement system before the model is ready. Keep the model integration modular.
Pro tip Separate model calls from the rest of the application so providers can be exchanged easily.
Watch out Avoid overbuilding customer-facing polish before the core workflow is validated.
- 4
Install a minimum viable model
Use the strongest current model to produce a working baseline. Treat weak outputs as diagnostic evidence rather than proof that the use case is impossible.
Pro tip Save representative inputs and outputs as a reusable evaluation set.
Watch out Do not expose unreliable results to customers without safeguards and human review.
- 5
Retest new model releases
Run each promising new model against the same evaluation set and quality thresholds. Compare performance, cost, consistency, and operational fit.
Pro tip Automate regression testing so model substitutions can be evaluated quickly.
Watch out A better benchmark score does not guarantee better performance in the specific workflow.
- 6
Activate when the threshold is crossed
Once a model meets the acceptance criteria, plug it into the prepared system and move the use case toward production. Continue monitoring because model behavior can change.
Pro tip Use staged rollout and human review before making the workflow fully autonomous.
Watch out Do not confuse a successful demonstration with dependable production performance.
In the wild
A software company expects AI to handle complex customer questions eventually, but current answers are inconsistent. It builds retrieval pipelines, permission controls, escalation logic, evaluations, and an internal interface while requiring agent approval. When a stronger model passes the stored test set, the company enables limited customer-facing responses without rebuilding the system.
→ The company converts a model release into a rapid, controlled product launch.
A marketing team creates recurring data collection, competitor definitions, reporting templates, confidence thresholds, and historical trend storage. The current model generates an internal draft only. New models are tested against known survey results until one becomes accurate enough for routine decision support.
→ Model improvements increase report quality immediately while preserving historical continuity.
Common mistakes
Building only for today's limitations
Restricting every use case to current capabilities leaves the team starting from zero when a stronger model arrives.
Creating a disposable model wrapper
A product with no proprietary workflow, data, distribution, or customer value can be erased when the provider adds the same feature.
Skipping baseline evaluation
Without stable test cases and thresholds, the team cannot tell whether a new model actually makes the intended workflow viable.
Is it for you?
Best for
Product, marketing, and innovation teams that believe model capabilities will soon make valuable but currently unreliable workflows viable.
Not ideal for
Teams without a durable use case or those building products whose only advantage is access to a third-party model.
From the transcript
“the other thing I would do is I would build use cases for where the models are going to be in 12 18 24 months…”
“I would build that all of the infrastructure to bring that use case to life and plug in the current model as a minimal viable…”
“as soon as that new model comes online or a new model comes online I could just plug it in and see the results get…”
From the episode
How This GPT-4 Prompt Is Breaking a $257 Billion Industry