Frontis-MA1, a 35B-parameter agent, scored 71.21% on MLE-Bench Lite, outperforming GPT-5.5 and Codex. The meta-evolution approach on OpenMLE suggests recursive self-improvement can rival trillion-scale pretraining.
Key facts
- 71.21% medal average on MLE-Bench Lite
- 35B parameter model beats GPT-5.5 and Codex
- Trained on OpenMLE, a recursive self-improvement stack
- Claim sourced from a single tweet, no paper linked
Frontis-MA1 (35B) beats GPT-5.5 + Codex on MLE-Bench Lite, posting a 71.21% medal average per @HuggingPapers. The model is a meta-evolution agent trained on OpenMLE, described as an open full-stack system for recursive self-improvement in machine learning engineering. The tweet provides no architectural details, training compute, or ablation data; the source is a single social post, so verification against a paper or benchmark leaderboard is pending.
Key Takeaways
- Frontis-MA1, a 35B meta-evolution agent, scored 71.21% on MLE-Bench Lite, beating GPT-5.5 and Codex.
- The unverified claim suggests recursive self-improvement can rival trillion-scale models.
What the 71.21% Means
MLE-Bench Lite is a reduced version of the Machine Learning Engineering benchmark, which tests agents on end-to-end ML tasks including data prep, model training, and evaluation. A 71.21% medal average on this subset puts Frontis-MA1 ahead of GPT-5.5 and Codex, per the claim. The significance is the parameter count: 35B versus trillion-scale rivals. If the result holds, it challenges the assumption that frontier capability requires frontier compute.
Recursive Self-Improvement as the Differentiator
The OpenMLE training pipeline reportedly uses meta-evolution — agents improving the systems that train them. This is distinct from the standard RLHF or supervised fine-tuning used on frontier models. The approach is not new in theory; Schmidhuber's 1987 Gödel machine and later self-referential architectures proposed it. Frontis-MA1 would be the first public evidence that the method produces competitive benchmark results at a fraction of the scale. The source does not disclose whether the 35B model is dense or MoE, nor the training data mix.
Skepticism Required

The claim comes from a single tweet, not a paper or reproducible leaderboard entry. No code, weights, or evaluation logs are linked. The benchmark itself is Lite, which may not reflect full MLE-Bench difficulty. Until the OpenMLE repository publishes the eval harness and model weights, the result should be treated as an unverified claim. The performance gap against trillion-scale rivals raises questions about parameter-count supremacy.
What to watch
Watch for the OpenMLE repository to publish the full evaluation harness and model weights. If the 71.21% reproduces independently, expect a wave of meta-evolution training runs. Also track whether MLE-Bench full results surface, as Lite subsets can flatter smaller models.









