Data Is the Currency Model
Trace AI product quality back to the scale and suitability of its data.
- Difficulty
- Moderate
- Time to result
- ~weeks to results
- Steps
- 6
- Confidence
- 98%
The Data Is the Currency Model treats usable data as a central input to AI product performance and competitive advantage. Begin with the output a product must generate, then trace backward to the scale, relevance, diversity, and quality of the data needed to produce it. Compare those requirements with the datasets each team can legally and practically access. In the episode, Adobe's restricted training set is contrasted with Midjourney's broader image access to explain a visible output gap despite Adobe's resources. The model predicts that a smaller team can beat a larger incumbent when its usable dataset is broader or better matched to the task. However, data advantage is not assessed in isolation: provenance, consent, licensing, legal exposure, model design, and evaluation quality remain essential constraints.
Origin
Extracted from Marketing Against The Grain's explanation of the output differences between Adobe Firefly and Midjourney.
Core principles
- 01Model quality depends on the data available to train and improve it.
- 02A smaller team with broader data can outperform a larger incumbent.
- 03Data restrictions can materially alter product output.
- 04Dataset advantage must be evaluated alongside legality, consent, and provenance.
How to run it
- 1
Specify the Output
Define what users consider a successful result and how it will be measured.
Pro tip Use representative benchmark tasks rather than subjective impressions alone.
Watch out A vague output target makes data requirements impossible to assess.
- 2
Map Required Data
Identify the examples, labels, styles, contexts, and edge cases the system must learn from.
Pro tip Separate data volume from relevance and coverage.
Watch out Large datasets can still be poorly matched to the desired task.
- 3
Audit Usable Data
Measure what data the organization can lawfully obtain, process, retain, and use for the intended purpose.
Pro tip Record provenance and license conditions at ingestion time.
Watch out Public availability does not automatically mean unrestricted training permission.
- 4
Compare Competitor Access
Estimate whether competitors have broader, cleaner, more proprietary, or more task-relevant datasets.
Pro tip Look for structural access advantages rather than temporary scraping volume.
Watch out External dataset estimates will often contain uncertainty.
- 5
Close the Data Gap
Develop lawful acquisition, licensing, partnership, contribution, or first-party data loops that improve the product.
Pro tip Prefer compounding first-party feedback loops when possible.
Watch out Do not pursue output quality by disregarding consent, copyright, or contractual boundaries.
- 6
Re-test Product Quality
Use the original benchmark to determine whether improved data actually changes output performance.
Pro tip Track gains by dataset revision so the cause remains attributable.
Watch out Do not assume every added dataset improves the model.
In the wild
The hosts compare image outputs and attribute part of the quality difference to Adobe Firefly being trained on Adobe Stock and open-licensed images while Midjourney draws from a broader image universe.
→ The comparison illustrates how usable dataset scope can outweigh company size or team resources.
A smaller software company trains and evaluates a support assistant using years of clean, permissioned tickets and resolution outcomes specific to its product. A larger generic competitor has more total text but less relevant support history.
→ The smaller company can outperform on its narrow domain through superior task-specific data.
Common mistakes
Equating Volume with Value
Data that is duplicated, noisy, biased, stale, or irrelevant may not improve the desired output.
Ignoring Provenance
A dataset advantage can become a liability when licensing, consent, or ownership cannot be demonstrated.
Discounting Other Inputs
Algorithms, compute, product design, evaluation, and distribution also affect whether data becomes a durable advantage.
Is it for you?
Best for
It is best for leaders evaluating AI vendors, product moats, and investments in proprietary or licensed datasets.
Not ideal for
It is not ideal for claiming that more data always beats better algorithms, cleaner data, stronger evaluation, or lawful governance.
From the transcript
“data is the currency of an AI world. I'm gonna say it again. Data is a currency of the AI world.”
“You can have an amazing model. You can have a great team like Adobe, but if you are using a much more limited data set,…”
“the difference in output between these is Mid Journey is basing these images off of all available images that they can get their hands on.”
From the episode
Midjourney V5 Update: Everything You Need To Know (#107)