Task-Based AI Model Portfolio
Assign each AI model a job, then test challengers against the incumbent.
- Difficulty
- Easy
- Time to result
- ~days to results
- Steps
- 5
- Confidence
- 94%
The framework treats AI models as a portfolio of specialist tools rather than forcing one model to handle every kind of work. Begin by separating recurring activities such as strategic thinking, personal assistance, research, writing, and coding. Assign an incumbent model to each activity based on demonstrated strengths, including integrations and privileged data sources. When a new model appears, do not migrate everything or chase benchmark headlines. Instead, split-test it against the incumbent using representative work from the relevant category. Compare the usefulness of the output, the amount of correction required, speed, and access to necessary context. Retain the current assignment unless the challenger creates a meaningful practical advantage. This produces a manageable AI stack that can evolve without requiring constant tool switching.
Origin
Extracted from Marketing Against the Grain during a discussion of how the hosts allocate Claude, Gemini, Grok, Google Deep Research, and OpenAI models across daily work.
Core principles
- 01Choose models by task rather than overall reputation.
- 02Reduce tool overload by maintaining a small working portfolio.
- 03Treat proprietary data access as a task-specific advantage.
- 04Test new models against the current tool on real work.
- 05Keep incumbents unless a challenger produces materially better results.
How to run it
- 1
Map Recurring Jobs
Divide your AI usage into stable job categories such as strategy, assistance, research, writing, and coding. Define the output you need from each category.
Pro tip Use categories based on actual weekly work rather than hypothetical use cases.
Watch out Do not create so many categories that maintaining the portfolio becomes another burden.
- 2
Assign Incumbents
Select one primary model for each job based on practical performance, integrations, and data access. Record the reason for each assignment.
Pro tip A model with access to the right private or real-time data may outperform a stronger general model.
Watch out Do not select solely from public benchmark rankings.
- 3
Run Real-Work Split Tests
Give a challenger and the incumbent the same representative task, context, and success criteria. Evaluate both outputs in the environment where they will actually be used.
Pro tip Use consequential but reversible work that you already understand well enough to judge.
Watch out A single unusually good response is not sufficient evidence of consistent superiority.
- 4
Compare Practical Utility
Assess accuracy, insight, speed, editing burden, integrations, and source coverage. Weight the criteria according to the job rather than using one universal score.
Pro tip Track whether the output moves the work forward, not merely whether it sounds impressive.
Watch out Do not confuse a model's conversational personality with reliable task performance.
- 5
Update Selectively
Replace an incumbent only where the challenger delivers a repeatable advantage. Leave the rest of the portfolio unchanged and revisit assignments periodically.
Pro tip A small number of intentional changes is easier to operationalize than a complete migration.
Watch out Constant switching destroys the familiarity and reusable workflows built around an incumbent.
In the wild
A marketing leader currently uses an OpenAI reasoning model for internal strategy. After Claude gains a thinking mode, the leader runs both models on the same HubSpot strategy problem, compares the resulting recommendations, and changes the strategy assignment only if Claude delivers a stronger thought-partner experience.
→ The new model is evaluated on relevant work without disrupting the rest of the AI stack.
A researcher uses Google and OpenAI deep-research tools for broad web research but reserves Grok for questions requiring current X data. The tools occupy different portfolio slots because their source access differs, even when their general capabilities overlap.
→ Research is routed according to source advantage instead of forcing one model to cover every information environment.
Common mistakes
Chasing Every Release
Trying every new model across every workflow creates cognitive overhead without proving that any change improves the work.
Using One Universal Ranking
A model that excels at coding may not be the best choice for research, assistance, or strategic reasoning.
Ignoring Data Access
Comparing models only on generated answers overlooks decisive advantages such as access to X, Maps reviews, or Google Drive.
Is it for you?
Best for
It is best for knowledge workers who use AI across strategy, research, writing, administration, and coding.
Not ideal for
It is not ideal for occasional users whose needs can be met by one general-purpose assistant.
From the transcript
“there's too much to try to keep up with.”
“For writing, I use Claude, and so I'm and so I'm going to continue. And then for coding, I use Claude.”
“the thing I'm going to try is it's Claude thinking model for the strategic stuff that I'm working on for HubSpot that I'm currently using,…”
From the episode
BREAKING: Claude 3.7 & Claude Code Just Dropped! (Massive AI Upgrade)