Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

An open laptop screen displays a web browser interface with a glowing AI agent icon hovering over a task completion…
AI ResearchScore: 90

Microsoft Fara1.5-27B Open-Source Agent Scores 72.3% on Web Tasks

Microsoft released Fara1.5-27B, a vision-only web browsing agent scoring 72.3% on Online-Mind2Web, open-source on Hugging Face.

·1d ago·2 min read··31 views·AI-Generated·Report error
Share:
What is Microsoft's Fara1.5-27B model?

Microsoft released Fara1.5-27B, a vision-only computer use agent for web browsing, scoring 72.3% on the Online-Mind2Web benchmark. The 27B-parameter model is open-source on Hugging Face.

TL;DR

Microsoft released Fara1.5-27B on Hugging Face · Vision-only computer use agent for web browsing · Scores 72.3% on Online-Mind2Web benchmark

Microsoft released Fara1.5-27B on Hugging Face, a vision-only computer use agent for web browsers. The 27B-parameter model scores 72.3% on the Online-Mind2Web benchmark.

Key facts

  • 27B parameter model released on Hugging Face
  • 72.3% on Online-Mind2Web benchmark
  • Vision-only, no HTML or DOM parsing
  • Open-source weights and inference code

Microsoft released Fara1.5-27B on Hugging Face, a vision-only computer use agent for web browsers. The 27B-parameter model scores 72.3% on the Online-Mind2Web benchmark. According to @HuggingPapers

Architecture and Approach

Fara1.5-27B operates solely on visual input — screenshots of the browser — without requiring structured HTML, accessibility trees, or DOM parsing. This mirrors the approach taken by earlier computer-use agents like Apple's Ferret-UI and Anthropic's computer-use mode for Claude, but with a focus on web browsing rather than general UI navigation.

The model is released under an open-source license on Hugging Face, with weights and inference code available. Microsoft did not disclose the training dataset size, compute budget, or training methodology beyond the model card.

Benchmark Context

Online-Mind2Web is an evaluation framework that measures end-to-end task completion on live websites, not static traces. A 72.3% success rate means the model correctly completes roughly 7 out of 10 web browsing tasks from scratch. For comparison, prior state-of-the-art results on this benchmark have hovered around 60-65% for similarly sized models, though direct apples-to-apples comparison is complicated by the dynamic nature of live web tasks.

The vision-only approach avoids the brittleness of HTML-parsing agents that break when websites change their DOM structure. However, it also introduces latency and resolution constraints from screenshot processing.

Implications for Open-Source Agents

Fara1.5-27B joins a growing ecosystem of open-source computer-use agents, including CogAgent (9B, from Tsinghua University) and WebGPT (OpenAI, not open-source). The release makes a competitive web-browsing agent available to researchers and developers without API costs or rate limits.

Microsoft has not published comparative results against closed-source alternatives like Claude's computer-use mode or GPT-4 with vision. The company's blog post says the model targets "research and prototyping use cases" rather than production deployment.

What to watch

Watch for independent replication of the 72.3% score by third parties, and for Microsoft to release a technical report detailing training data and compute. Also track whether the model generalizes to non-English websites or dynamic single-page apps.

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from multiple verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

Fara1.5-27B represents a meaningful step in open-source computer-use agents, but the 72.3% score needs independent verification. The vision-only approach trades off the structural precision of DOM parsing for robustness against website changes — a bet that may pay off for dynamic content but introduces latency from screenshot processing. Microsoft's silence on training data and compute is a red flag. Without reproducibility details, the model's true capabilities remain uncertain. The release also lacks comparison to closed-source alternatives, making it hard to gauge where this sits relative to production systems. The open-source nature is the key differentiator. Researchers can now fine-tune and adapt a competitive web agent without API costs, potentially accelerating progress in areas like web automation, accessibility tools, and AI-powered testing.
This story is part of
Claude Code's Campus Conquest Flips Anthropic's Talent Pipeline, Leaving Google's Academic Edge in Doubt
Viral adoption at MIT and Stanford transforms Claude Code from product into recruiting funnel, threatening Google's long-held research talent dominance
Compare side-by-side
Microsoft vs Hugging Face

Mentioned in this article

Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in AI Research

View all