Ethan Mollick, Wharton professor, argues negative AI findings on GPT-4 age poorly, while positive ones persist. He outlines five publishable frames for claims about older models in a recent post.
Key facts
- Mollick: positive AI findings persist across generations
- Negative claims require explicit model naming or trends
- GPT-5.6 Sol cited as example reasoning model
- Reproducibility turns negative results into benchmarks
- Human-centric studies remain durable regardless of model
In a recent post on X, Ethan Mollick, a Wharton professor known for AI commentary, argues that research on older models like GPT-4 is not inherently flawed—but negative claims demand far more care than positive ones.
Why Positive Claims Persist
Mollick's core observation: "Generally, once AI has gained an ability, it does not regress in future generations." Capability thresholds—like "good as a human"—established with GPT-4 remain valid as a floor. Papers showing a minimum impact or effect hold up even if the model is superseded.
The Negative Claim Problem
For negative findings—bias, errors, or task failures—the logic breaks. "You cannot claim that because GPT-4 fails at something, that AI is bad at that task, because current or future models may do it." Mollick offers five frames to keep such work publishable:
- Explicit model naming: "GPT-4 could not do X." Rarely interesting alone, but with reproduction details it becomes a benchmark.
- Trend measurement: Compare GPT-4 to GPT-5 to GPT-5.6 Sol, including at least one reasoning model, to chart ability trajectories.
- Grounded limitations: Argue a natural flaw prevents the task, supported by evidence.
- Moderator/mediator focus: Show how prompting, context, or social factors affect GPT-4's performance—a forward-looking concern.
- Human-centric analysis: Study reactions, dangers, or advantages, independent of model version.
Mollick concludes: "generally negative capability claims have been much less durable than positive ones." The takeaway for researchers: positive results are safe; negative ones need explicit framing or a trend line to stay relevant.
What to watch
Watch for a surge in trend-based studies comparing GPT-4 to GPT-5.6 Sol, particularly in bias and safety literature. If top venues like NeurIPS or ICML start rejecting single-model negative claims without reasoning-model baselines, Mollick's framing becomes de facto policy.







