OSReward
product→ stable
OSReward is a human-gold benchmark developed to evaluate VLM judges scoring computer-use agents, providing a reference standard for assessing agent performance. It serves as a critical tool for measuring and comparing the accuracy of vision-language
1Total Mentions
+0.60Sentiment (Very Positive)
+1.2%Velocity (7d)
First seen: Aug 7, 2026Last active: 1d ago
Signal Radar
Five-axis snapshot of this entity's footprint
Loading radar…
Mentions × Lab Attention
Weekly mentions (solid) and average article relevance (dotted)
mentionsrelevance
Loading timeline…
Timeline
1- Product LaunchAug 7, 2026
OSReward benchmark and OS-Shepherd reward models announced, claiming 30-60x lower cost than commercial judges.
View source
Relationships
1Uses
Predictions
No predictions linked to this entity.
AI Discoveries
No AI agent discoveries for this entity.
Sentiment History
Positive sentiment
Negative sentiment
Range: -1 to +1
| Week | Avg Sentiment | Mentions |
|---|---|---|
| 2026-W32 | 0.60 | 1 |