MMarketing Against The Grain
← All episodes
15 July 2025

Is Grok 4 Worth $30/mo? (Full Test)

2Frameworks
9Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 1

Myth Buster01:00

Why Benchmark Dominance Does Not Guarantee Better Marketing

Grok 4 leads respected reasoning benchmarks, but the hosts caution that benchmark performance can be engineered or disconnected from practical quality. The episode therefore tests marketing and knowledge-work tasks rather than accepting benchmark rankings as a verdict.

  • ARC-AGI is presented as a respected independent benchmark.
  • Models can perform well on benchmarks while remaining weak at particular tasks.
  • Real-world applications are necessary to judge practical usefulness.

But apparently you can for sure uh appear pretty well in benchmarks while still not being great at certain tasks.

Kieran Flanagan · 01:30
#grok-4#benchmarks#ai-evaluation

Hot Take· 2

Hot Take03:00

AI Models Are Still Improving Too Fast to Settle on One Winner

The hosts argue that model progress has not reached a plateau and expect rapid competitive responses from OpenAI and Google. Grok 4's lead is framed as temporary evidence of a continuing innovation cycle rather than a permanent shift in the market.

  • Model capability is still improving materially.
  • Competitive labs are unlikely to leave one model in front for long.
  • Users should expect frequent changes in the preferred model.
  • The current model landscape may look very different within a few years.

AI innovation curve is still really steep.

Kip Bodnar · 03:00

And so you're going to continue to see this wave of innovation be very fast and very steep.

Kip Bodnar · 03:00
#ai-models#innovation#competition
Hot Take14:30

Grok 4's Sam Parr-Style Newsletter Is Only Fine

Grok analyzes recent posts and produces a newsletter inspired by Sam Parr's style, but the result does not convincingly reproduce his writing. Kieran finds the creative output acceptable rather than exceptional and says the same prompt performed much better with OpenAI's o3.

  • The prompt analyzes recent long-form posts and engagement.
  • Grok incorporates material connected to Sam Parr's podcast.
  • The generated writing only partially resembles the target style.
  • OpenAI o3 reportedly produces a much stronger result on the same prompt.

It's a little bit like Sam knowing Sam's writing. It's not exactly like him.

Kieran Flanagan · 15:30

the output was much much better than grock 4.

Kieran Flanagan · 16:00
#copywriting#creative-ai#grok-4#o3

Explainer· 2

Explainer02:00

How Grok 4 Heavy Uses Competing Agent Swarms

The premium Heavy model reportedly sends multiple agents to solve the same task and then selects the strongest result. The approach could improve difficult problem solving, but it is slower and tied to a substantially more expensive subscription.

  • Multiple agents attempt the same task.
  • The system compares agent solutions before returning an answer.
  • The Heavy tier reportedly costs about $300 per month.
  • Complex runs can take considerable time.

It has a whole team of agents go and try to solve a task and then it basically uh looks to see what agent has…

Kieran Flanagan · 02:00

And uh it's it's pretty timeconuming

Kieran Flanagan · 02:30
#grok-4-heavy#agent-swarms#reasoning
Explainer16:30

Grok 4 Builds a Full Campaign but Struggles With Creativity

A fictional AI note-taking startup is used to test chained reasoning across personas, competition, risks, validation, messaging, channel mix, and self-assessment. Grok completes the structure and identifies plausible risks, but its differentiation and campaign copy are weakened by limited company context and mediocre creativity.

  • The task combines reasoning, logic, math, and multiple chained outputs.
  • Grok produces personas, competitive analysis, risks, and a one-week plan.
  • Missing company context forces the model to invent differentiators.
  • The campaign's structure is stronger than its creative messaging.

There's reasoning. There's a lot of math and logic. There's a bunch of different steps which is like multi-ch kind of multi-chaining together tasks.

Kieran Flanagan · 18:30

Yeah, the actual creativity behind this is not very good.

Kieran Flanagan · 20:00
#campaign-planning#marketing-ai#reasoning

Q&A· 1

Q&A09:00

Can You Spot Grok 4's Image Quality Upgrade?

A blind comparison of Vermont summer images shows Grok 4 producing greater sharpness, vibrancy, and fine detail than Grok 3. The composition is not radically different, but the newer output appears more refined.

  • The Grok 3 image looks more cartoon-like.
  • Grok 4 shows stronger fine detail and sharpness.
  • Color vibrancy is noticeably improved.
  • The gain is refinement rather than a fundamentally different image.

I think the first one's Grock 3 cuz it's a little more cartoony.

Kieran Flanagan · 09:30

the detail sharpness is significantly better.

Kip Bodnar · 10:00
#image-generation#grok-4#multimodal

Takeaway· 3

Takeaway11:00

Grok 4 Is More Compelling for Solo Users Than Enterprises

The hosts distinguish individual value from organizational readiness. Grok 4 offers strong model quality for its price, but it lacks the mature business controls and scaling features associated with ChatGPT or Anthropic products.

  • Grok is positioned more strongly for individual users.
  • Business controls and scalability trail enterprise-oriented competitors.
  • The $30 tier offers competitive model quality.
  • The Heavy tier is necessary for fuller access to premium reasoning.

They are like much more solo user than business user models.

Kip Bodnar · 11:00

I think the $30 Grock is now pretty close to Chat GPT's $200 a month subscription.

Kip Bodnar · 11:30
#pricing#enterprise-ai#grok-4
Takeaway28:00

Poor Context, Not Just the Model, Limits Marketing Output

Kieran acknowledges that several tests omitted essential inputs, including a real company profile, a detailed content ICP, and examples of successful work. Those omissions make generic or invented recommendations more likely, so the episode's results are partly a warning about prompt context as well as model quality.

  • Fake brands leave the model without factual differentiation.
  • Audience profiles should accompany content-generation requests.
  • Successful prior examples help guide style and quality.
  • Model comparisons are less meaningful when prompts lack production context.

I really should have a whole profile from my ideal customer profile.

Kieran Flanagan · 28:30

I should have examples of content that we've created that have done really well.

Kieran Flanagan · 28:30
#prompting#context#content-quality
Takeaway28:30

The Verdict: Use Grok 4 for X Data, Not as an o3 Replacement

After weekend testing and several marketing tasks, Kieran concludes that Grok 4 is competent but not transformative for general marketing and knowledge work. Its distinctive advantage is access to X data, while o3 remains his preferred option for many broader tasks.

  • Grok 4 is judged capable rather than exceptional.
  • Benchmark leadership does not translate into a clear marketing breakthrough.
  • Access to X data creates several specialized use cases.
  • The host does not plan to move most work away from o3.

It's not incredible. I have not found it to be incredible.

Kieran Flanagan · 28:30

I'm going to use it for a bunch of things where I need access to the X data, the Twitter data.

Kieran Flanagan · 29:30
#grok-4#o3#ai-tools#verdict