MMarketing Against The Grain
← All episodes
20 February 2025

GROK 3 vs GPT-4: The AI War Just Got Real [First Look]

7Frameworks
14Insights

Listen

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 1

Myth Buster08:00

Getting Better AI Output Does Not Mean You Are Getting Better

The hosts warn that offloading writing or coding to AI can produce useful results without building the user’s underlying competence. Kieran describes deliberately studying Cursor so that his skills improve alongside his use of AI-assisted development tools.

  • AI can reduce expert work to a few text-box instructions
  • Successful output may conceal a lack of understanding
  • Heavy offloading can weaken writing or coding ability
  • Users should deliberately learn alongside AI

you can just lose the art of writing because you're offloading a lot of that to an AI assistant.

Kieran · 08:30

I'm trying to like really make sure that I'm getting better along with my AI usage versus me just offloading the things that I want…

Kieran · 09:00
#domain-expertise#ai-coding#learning#skill-development

Hot Take· 3

Hot Take03:00

Why Grok’s Enterprise Push Makes It a Serious OpenAI Rival

The discussion moves beyond consumer chat to Grok’s availability through Palantir and its apparent enterprise ambitions. The hosts argue that Elon Musk’s capital, execution record, and rivalry with Sam Altman make xAI a particularly formidable competitor.

  • Palantir introduced Grok into enterprise environments
  • xAI is targeting more than consumer use cases
  • Elon Musk has unusual access to capital and infrastructure
  • Personal rivalry may intensify competition with OpenAI

Palantur has brought Grok to the enterprise now, officially available.

Kieran · 03:00

I do think that this is problematic for open AI because you've got a pretty incredible competitor in Elon, which considering this team started a…

Kieran · 03:30
#enterprise-ai#openai#palantir#competition
Hot Take06:00

When Intelligence Is Free, Motivation Becomes the Advantage

As advanced models approach or surpass expert-level performance, access to intelligence becomes less scarce. The hosts contend that ideas, execution, behavior, emotional capability, and sustained motivation will increasingly determine who creates value.

  • AI is democratizing access to expert-level intelligence
  • Scarcity once made intelligence and education more economically valuable
  • Execution and behavior become stronger differentiators
  • Long-term motivation determines whether intelligence is applied

It shows you how less important intelligence is than what everybody thought.

Kipp · 06:30

differentiation is ideas, execution, behavior, motivation.

Kieran · 08:00
#intelligence#motivation#execution#future-of-work
Hot Take26:30

Enterprise AI Was Waiting for Access to Reasoning Models

The closing discussion argues that limited access—not merely model quality—has constrained enterprise adoption of advanced reasoning. Once reasoning models became available within the hosts’ enterprise environment, Kieran experienced an orders-of-magnitude improvement in using AI as a problem-solving partner.

  • Enterprise reasoning access had been heavily API-dependent
  • Seat-based access to reasoning models arrived only recently
  • Direct enterprise access unlocks internal-data use cases
  • Reasoning materially improves collaborative problem-solving

reasoning in companies is still very limited to the API.

Kipp · 26:30

The difference is orders of magnitude difference and me being able to work through problems as a thought partner with AI.

Kieran · 27:00
#enterprise-ai#api#reasoning-models#adoption

Explainer· 2

Explainer00:30

The Compute Behind Grok 3’s Benchmark Performance

Grok 3 debuted as the first model reported to exceed 1400 in the Chatbot Arena while outperforming leading reasoning models from OpenAI and Google. The hosts connect that performance to xAI’s rapid construction of a 200,000-GPU cluster in Memphis.

  • Grok 3 reportedly exceeded 1400 in Chatbot Arena testing
  • The model outperformed leading OpenAI and Google reasoning models
  • xAI assembled 200,000 GPUs in 214 days
  • The Memphis facility required an extraordinary amount of temporary cooling

it is the first model ever to score over 1400 on the chatbot arena and out perforbs the best reasoning models from OpenAI and Google.

Kipp · 00:30

they went from nothing to 200,000 GPUs in 214 days.

Kipp · 01:30
#grok-3#benchmarks#gpu-infrastructure#xai
Explainer09:30

Grok 3 Replaces Model Menus With Simple Capability Toggles

Grok’s interface minimizes model-selection complexity by presenting Grok 2 and Grok 3 while exposing research and reasoning as straightforward options. The hosts find this more approachable than products that require consumers to distinguish among numerous specialized model variants.

  • Users primarily choose between Grok 2 and Grok 3
  • Deep research and reasoning are integrated capabilities
  • The interface avoids separate model variants for each mode
  • Simpler controls improve the consumer experience

Just one model.

Kipp · 10:00

The user experience is a little simpler and more consumer-friendly here than I think on some of the other models today.

Kipp · 10:30
#grok-3#user-experience#reasoning-models#deep-research

Story· 2

Story17:00

Five Years of Internal Data Turned AI Into a Strategic Colleague

Kieran describes loading five years of company research, metrics, strategic documents, goals, hypotheses, and results into a reasoning model. Instead of requesting a final answer, he spent three hours interrogating trends and problems, retrieving supporting evidence, and refining his thinking as if working with a colleague.

  • Rich internal context dramatically improved the model’s usefulness
  • The source material covered five years of strategy and results
  • Iterative questioning exposed important trends and problems
  • The model retrieved relevant lines from many documents
  • The session functioned as collaborative problem-solving rather than answer generation

But when you actually start to give them internal context, they're amazing.

Kieran · 17:00

I spent three hours this morning working with it like a colleague.

Kieran · 17:30
#internal-data#strategic-thinking#reasoning-models#enterprise-ai
Story21:30

Grok Found Seasonality and Personas in Synthetic Campaign Data

A marketing test asks Grok to analyze synthetic campaign history for a hypothetical medical-software company and build a future campaign plan. The model identifies seasonality, customer personas, healthcare events, policy changes, and industry shifts while producing a cleanly structured response.

  • The test used synthetic historical marketing data
  • Grok proposed a lead-generation campaign for private clinics
  • The model detected deliberately embedded seasonal factors
  • It identified personas and external healthcare trends
  • Its output closely resembled the earlier DeepSeek result

It actually built me out a marketing annual plan for the channels based upon this data in a way that I think it would have…

Kieran · 22:00

it got that, it got the different personas.

Kieran · 22:30
#marketing-analysis#campaign-planning#synthetic-data#grok-3

Q&A· 1

Q&A14:00

Can Grok 3 Produce a Strong YouTube Growth Strategy?

The hosts test Grok with raw YouTube performance data previously given to OpenAI reasoning models. Grok’s non-thinking response is judged basic, while its reasoning response improves to roughly O1-level quality but remains more conventional than the strongest O1 Pro strategy work.

  • The same prompt and dataset were used for model comparison
  • Grok’s non-thinking response was generic
  • Reasoning surfaced a viral Cody Sanchez short as an outlier
  • The improved recommendations covered topics, thumbnails, short-form content, collaborators, and watch time
  • Kieran still preferred O1 Pro for deeper strategy

this version I'm showing first is the non-thinking version.

Kipp · 15:00

But the thing I would say is the O1 Pro model is really great at strategy.

Kieran · 16:00
#youtube-growth#marketing-strategy#model-comparison#grok-3

Tool· 2

Tool10:30

Grok Deep Research Turns 45 Sources Into an Evidence Table

A test on red-light therapy shows Grok searching 45 web pages, displaying its research stages, and producing a structured report. The output separates health conditions, possible benefits, evidence strength, safety information, and citations without requiring an elaborate initial prompt.

  • The test searched 45 web pages
  • Grok exposed the stages of its research process
  • The report summarized benefits, safety, usage, and citations
  • A table connected conditions to benefits and evidence strength
  • The model needed fewer clarifying questions than ChatGPT often does

So it searched 45 web pages.

Kipp · 11:00

it put a table together, Kieran, of the condition you might have, the benefit, and whether there's like real evidence or not, right?

Kipp · 13:00
#deep-research#red-light-therapy#evidence#grok-3
Tool19:00

Give Colleagues Both Human and AI Feedback

Kieran explains that he now supplements his personal feedback with an assessment from the model best suited to the problem. He clearly distinguishes where he agrees or disagrees with the model, giving recipients another analytical perspective without presenting AI output as unquestioned truth.

  • Select a model suited to the specific problem
  • Provide personal feedback first
  • Include the model’s separate assessment
  • State where human judgment agrees or disagrees

I give them my feedback and I tell them I also give them the feedback from the model I think is most appropriate for that…

Kieran · 19:30

And I tell them what I agree with or disagree with that the model says

Kieran · 19:30
#feedback#management#ai-assistance#decision-support

Takeaway· 3

Takeaway05:00

Why AI Benchmarks Cannot Pick the Best Model for You

Benchmark scores establish a model’s general capability, but they do not determine whether it will excel at a particular user’s work. The hosts recommend judging models through the real use cases and inputs that matter to the individual or organization.

  • Benchmarks provide useful baseline evidence
  • Model quality varies by use case
  • Practical testing matters more to users than leaderboard position
  • The best model is task-dependent

the proof is in the use cases that you want to use it for.

Kieran · 05:00

how important they are to the user, it depends upon the use case you're trying to do it for, and that's how you'll feel about…

Kieran · 05:00
#ai-benchmarks#model-selection#use-cases
Takeaway18:30

Use Reasoning Models to Attack Your Own Strategy

The hosts identify critical challenge and counterargument generation as an underused application of advanced reasoning models. Strategic professionals can ask a model to steel-man opposing positions, construct the strongest bullish case, or identify holes in their assumptions before committing to a decision.

  • Reasoning models can steel-man competing arguments
  • They can develop the strongest case for a position
  • They are useful for exposing gaps in strategic thinking
  • Frequent model dialogue is becoming essential in strategic roles

you can basically ask it to steel man an argument.

Kipp · 18:30

What I ask it to do is like poke holes in my thinking here

Kieran · 19:00
#critical-thinking#strategy#reasoning-models#decision-making
Takeaway23:00

Grok 3’s Early Edge: Speed, Structure, and Less Prompting

The hosts’ initial verdict is that Grok 3 produces well-structured output quickly and performs particularly well for users who do not write highly detailed prompts. They position it as a cheaper, faster option for repetitive reasoning work, while reserving judgment on specialized tasks such as writing.

  • Grok structures responses well without extensive instructions
  • Reasoning outputs arrive much faster than some competing models
  • It suits users who are not advanced prompt writers
  • The hosts recommend it for fast repetitive reasoning tasks
  • Further testing is needed for specialized use cases

The two things I would say I like most about it are that and that it is frickin' fast.

Kipp · 23:00

Grok is really good if you are not a detailed prompter. And if you're in a hurry, Grok is fast.

Kipp · 25:30
#grok-3#prompting#speed#model-selection