MMarketing Against The Grain
← All frameworks
Productivity

Task-Based AI Model Portfolio

Assign each AI model a job, then test challengers against the incumbent.

Difficulty
Easy
Time to result
~days to results
Steps
5
Confidence
94%

The framework treats AI models as a portfolio of specialist tools rather than forcing one model to handle every kind of work. Begin by separating recurring activities such as strategic thinking, personal assistance, research, writing, and coding. Assign an incumbent model to each activity based on demonstrated strengths, including integrations and privileged data sources. When a new model appears, do not migrate everything or chase benchmark headlines. Instead, split-test it against the incumbent using representative work from the relevant category. Compare the usefulness of the output, the amount of correction required, speed, and access to necessary context. Retain the current assignment unless the challenger creates a meaningful practical advantage. This produces a manageable AI stack that can evolve without requiring constant tool switching.

Origin

Extracted from Marketing Against the Grain during a discussion of how the hosts allocate Claude, Gemini, Grok, Google Deep Research, and OpenAI models across daily work.

Core principles

  • 01Choose models by task rather than overall reputation.
  • 02Reduce tool overload by maintaining a small working portfolio.
  • 03Treat proprietary data access as a task-specific advantage.
  • 04Test new models against the current tool on real work.
  • 05Keep incumbents unless a challenger produces materially better results.

How to run it

  1. 1

    Map Recurring Jobs

    Divide your AI usage into stable job categories such as strategy, assistance, research, writing, and coding. Define the output you need from each category.

    Pro tip Use categories based on actual weekly work rather than hypothetical use cases.

    Watch out Do not create so many categories that maintaining the portfolio becomes another burden.

  2. 2

    Assign Incumbents

    Select one primary model for each job based on practical performance, integrations, and data access. Record the reason for each assignment.

    Pro tip A model with access to the right private or real-time data may outperform a stronger general model.

    Watch out Do not select solely from public benchmark rankings.

  3. 3

    Run Real-Work Split Tests

    Give a challenger and the incumbent the same representative task, context, and success criteria. Evaluate both outputs in the environment where they will actually be used.

    Pro tip Use consequential but reversible work that you already understand well enough to judge.

    Watch out A single unusually good response is not sufficient evidence of consistent superiority.

  4. 4

    Compare Practical Utility

    Assess accuracy, insight, speed, editing burden, integrations, and source coverage. Weight the criteria according to the job rather than using one universal score.

    Pro tip Track whether the output moves the work forward, not merely whether it sounds impressive.

    Watch out Do not confuse a model's conversational personality with reliable task performance.

  5. 5

    Update Selectively

    Replace an incumbent only where the challenger delivers a repeatable advantage. Leave the rest of the portfolio unchanged and revisit assignments periodically.

    Pro tip A small number of intentional changes is easier to operationalize than a complete migration.

    Watch out Constant switching destroys the familiarity and reusable workflows built around an incumbent.

In the wild

Testing Claude for Strategy Work

A marketing leader currently uses an OpenAI reasoning model for internal strategy. After Claude gains a thinking mode, the leader runs both models on the same HubSpot strategy problem, compares the resulting recommendations, and changes the strategy assignment only if Claude delivers a stronger thought-partner experience.

The new model is evaluated on relevant work without disrupting the rest of the AI stack.

Separating General and Social Research

A researcher uses Google and OpenAI deep-research tools for broad web research but reserves Grok for questions requiring current X data. The tools occupy different portfolio slots because their source access differs, even when their general capabilities overlap.

Research is routed according to source advantage instead of forcing one model to cover every information environment.

Common mistakes

Chasing Every Release

Trying every new model across every workflow creates cognitive overhead without proving that any change improves the work.

Using One Universal Ranking

A model that excels at coding may not be the best choice for research, assistance, or strategic reasoning.

Ignoring Data Access

Comparing models only on generated answers overlooks decisive advantages such as access to X, Maps reviews, or Google Drive.

Is it for you?

Best for

It is best for knowledge workers who use AI across strategy, research, writing, administration, and coding.

Not ideal for

It is not ideal for occasional users whose needs can be met by one general-purpose assistant.

From the transcript

there's too much to try to keep up with.

Kieran Flanagan · 18:30

For writing, I use Claude, and so I'm and so I'm going to continue. And then for coding, I use Claude.

Kieran Flanagan · 19:30

the thing I'm going to try is it's Claude thinking model for the strategic stuff that I'm working on for HubSpot that I'm currently using,…

Kieran Flanagan · 19:30

From the episode

BREAKING: Claude 3.7 & Claude Code Just Dropped! (Massive AI Upgrade)