Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Dashboard showing benchmark scores comparing Frontis-MA1 at 71.21% against GPT-5.5 and Codex, with a line chart…
AI ResearchScore: 92

Frontis-MA1 35B Beats GPT-5.5 on MLE-Bench Lite

Frontis-MA1, a 35B meta-evolution agent, scored 71.21% on MLE-Bench Lite, beating GPT-5.5 and Codex. The unverified claim suggests recursive self-improvement can rival trillion-scale models.

·1d ago·3 min read··20 views·AI-Generated·Report error
Share:
How did the 35B Frontis-MA1 model beat GPT-5.5 on MLE-Bench Lite?

Frontis-MA1, a 35B meta-evolution agent trained on OpenMLE, scored 71.21% medal average on MLE-Bench Lite, beating GPT-5.5 and Codex. The result approaches trillion-scale system performance, suggesting recursive self-improvement can close the parameter-count gap.

TL;DR

35B model scores 71.21% on MLE-Bench Lite · Beats trillion-scale GPT-5.5 and Codex · Trained via meta-evolution on OpenMLE stack

Frontis-MA1, a 35B-parameter agent, scored 71.21% on MLE-Bench Lite, outperforming GPT-5.5 and Codex. The meta-evolution approach on OpenMLE suggests recursive self-improvement can rival trillion-scale pretraining.

Key facts

  • 71.21% medal average on MLE-Bench Lite
  • 35B parameter model beats GPT-5.5 and Codex
  • Trained on OpenMLE, a recursive self-improvement stack
  • Claim sourced from a single tweet, no paper linked

Frontis-MA1 (35B) beats GPT-5.5 + Codex on MLE-Bench Lite, posting a 71.21% medal average per @HuggingPapers. The model is a meta-evolution agent trained on OpenMLE, described as an open full-stack system for recursive self-improvement in machine learning engineering. The tweet provides no architectural details, training compute, or ablation data; the source is a single social post, so verification against a paper or benchmark leaderboard is pending.

Key Takeaways

  • Frontis-MA1, a 35B meta-evolution agent, scored 71.21% on MLE-Bench Lite, beating GPT-5.5 and Codex.
  • The unverified claim suggests recursive self-improvement can rival trillion-scale models.

What the 71.21% Means

MLE-Bench Lite is a reduced version of the Machine Learning Engineering benchmark, which tests agents on end-to-end ML tasks including data prep, model training, and evaluation. A 71.21% medal average on this subset puts Frontis-MA1 ahead of GPT-5.5 and Codex, per the claim. The significance is the parameter count: 35B versus trillion-scale rivals. If the result holds, it challenges the assumption that frontier capability requires frontier compute.

Recursive Self-Improvement as the Differentiator

The OpenMLE training pipeline reportedly uses meta-evolution — agents improving the systems that train them. This is distinct from the standard RLHF or supervised fine-tuning used on frontier models. The approach is not new in theory; Schmidhuber's 1987 Gödel machine and later self-referential architectures proposed it. Frontis-MA1 would be the first public evidence that the method produces competitive benchmark results at a fraction of the scale. The source does not disclose whether the 35B model is dense or MoE, nor the training data mix.

Skepticism Required

another important benchmark: gpt-5.5 sets a new SOTA o…

The claim comes from a single tweet, not a paper or reproducible leaderboard entry. No code, weights, or evaluation logs are linked. The benchmark itself is Lite, which may not reflect full MLE-Bench difficulty. Until the OpenMLE repository publishes the eval harness and model weights, the result should be treated as an unverified claim. The performance gap against trillion-scale rivals raises questions about parameter-count supremacy.

What to watch

Watch for the OpenMLE repository to publish the full evaluation harness and model weights. If the 71.21% reproduces independently, expect a wave of meta-evolution training runs. Also track whether MLE-Bench full results surface, as Lite subsets can flatter smaller models.

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from multiple verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The claim, if verified, would be a structural challenge to the scaling-law consensus. The AI field has spent three years assuming capability scales with parameters and data. A 35B model beating trillion-scale rivals on a meaningful benchmark would suggest that training methodology — specifically recursive self-improvement — can substitute for raw compute. This echoes the 2020 discovery that larger learning rates and better optimizers could close gaps previously attributed to model size. The meta-evolution framing is the key differentiator. Standard fine-tuning optimizes weights against a fixed objective. Meta-evolution optimizes the training system itself, allowing the agent to propose changes to its own training pipeline. This is computationally expensive per iteration but could produce compounding gains that static pipelines cannot. The 71.21% figure, if real, would be the first public evidence that this compounding effect produces frontier-competitive results. The lack of a paper is the critical weakness. One tweet with a benchmark number and no methodology is not science. The OpenMLE project needs to publish the eval harness, model weights, and training logs. Until then, the result sits in the same category as other unverified benchmark claims that populate X. The parameter-count question remains open, but the burden of proof is on the claim's authors.
This story is part of
The Protocol Schism: Anthropic's MCP Stack vs. OpenAI's Agent Lock-In
How a developer convention is splitting AI into two incompatible ecosystems, with Meta and Google caught in the middle
Compare side-by-side
Frontis-MA1 vs GPT-5.5
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in AI Research

View all