test run
30 articles about test run in AI news
Beijing Humanoid Robots Tested in Half-Marathon for Stability, Endurance
Humanoid robots in Beijing underwent a half-marathon test run, demonstrating sustained running speeds that challenge their dynamic stability and energy efficiency. This is a significant endurance test for real-world deployment.
How to Wire an Open-Source Coding Agent into Your Workflow With MCP — and
Wire openai/codex into your local workflow via MCP with strict boundaries (no test edits, no commits). Verify its patch with an independent test runner and git diff before human review. Works: 0/2 → 2/2 tests in 58 seconds.
Anthropic Hits $47B Run-Rate Revenue, Claims Fastest Organic Growth Ever
Anthropic's run-rate revenue hits $47B, up 57% from $30B, per Axios CEO Jim VandeHei—claimed as fastest organic growth ever.
ExSpec: Run Gherkin Tests in Real Browsers with Claude Code—No Step Definitions Required
ExSpec lets you write plain-text Gherkin specs and have Claude Code execute them in a real browser, eliminating brittle step definitions and glue code.
Dusk MCP: Stop Having Your AI Agent Guess Its Way Through Flutter Testing
Dusk MCP lets Claude Code drive a running Flutter app via the Semantics tree—no test files, no screenshot guessing. The 6-step actionability gate prevents flaky taps.
Beijing Humanoid Robot Half Marathon Tests 40% Autonomous Teams
A night-time half-marathon test for humanoid robots in Beijing revealed approximately 40% of participating teams were running fully autonomous systems, a key benchmark for real-world robotic mobility.
Keygraph's Shannon AI Pentester Hits 96.15% on XBOW, Finds Real Exploits
Keygraph released Shannon, a fully autonomous AI pentester that hunts real exploits in source code with a 96.15% success rate on the hint-free XBOW Benchmark. It runs a full test in about an hour for roughly $50 using Claude Sonnet.
Claude Code Hooks: How to Auto-Format, Lint, and Test on Every Save
Configure hooks in .claude/settings.json to run prettier, eslint, and tests automatically, ensuring clean code without manual intervention.
CMU Research Identifies 'Biggest Unlock' for Coding Agents: Strategic Test Execution
New research from Carnegie Mellon University suggests the key advancement for AI coding agents lies not in raw code generation, but in developing strategies for how to run and interpret tests. This shifts focus from LLM capability to agentic reasoning.
Tessera Launches Open-Source Framework for 32 OWASP AI Security Tests, Benchmarks GPT-4o, Claude, Gemini, Llama 3
Tessera introduces the first open-source framework to run all 32 OWASP AI security tests against any model with one CLI command. It provides benchmark results for GPT-4o, Claude, Gemini, Llama 3, and Mistral across 21 model-specific security tests.
Fortress Framework Prunes Unstable Features, Boosts Rec Stability by CV
Fortress prunes temporally unstable features in rec models via historical snapshots, improving CV and PR-AUC in offline tests.
Beijing Humanoid Robot Half-Marathon Test Ends in 'Mechanical Carnage'
A pre-race test for Beijing's 2026 humanoid half-marathon showed robots collapsing, exposing endurance gaps ahead of the 21.1km April event.
Anthropic Hits $65B Run Rate, Adding $18B in 2 Months
Anthropic's run rate hit $65B, adding $18B in two months, ahead of an October IPO targeting $2T+ valuation.
GLM-5.2 Nears Claude Opus 4.8 on Cyber-Defense Tests — Here's Why Claude
GLM-5.2 rivals Mythos 5 on cyber-defense. For Claude Code users: expect price cuts, better Opus 4.8 security, and new MCP options. Test your workflows with /model.
Spotify Engineers: LLM A/B Tests Recover Only 39% of Human Treatment Effects
Spotify Engineering tested LLM-based A/B testing on the Upworthy dataset, finding raw predictions recover only 39% of human treatment effects. The bias is systematic, and calibration works only under unverifiable assumptions.
Anthropic's unreleased model pushes Riemann bound, tests 650 ideas
Anthropic's unreleased model raised the lower bound for the Riemann hypothesis, testing 650 ideas with 60 subagents, confirmed by mathematicians and Lean. This signals AI's growing role in mathematical discovery.
selenium-mcp: Let Claude Code Write Real Selenium Tests by Seeing the Page
Install selenium-mcp to give Claude Code a real browser. It writes standard Selenium tests with accurate locators from live page inspection, integrating into your existing Java project.
How to Test and Debug MCP Servers for Claude Code: A Production Guide
Unit-test tool logic, mock external services, and add integration tests for MCP servers. Claude Code's MCP integration demands observability beyond local demos.
Brunello Cucinelli's Callimacus AI platform becomes a 7-figure software
Brunello Cucinelli's AI platform Callimacus became an independent software company, generating seven-figure revenue and attracting a Salesforce investment. The platform assembles personalized webpages in real time and is being pitched to brands and retailers beyond luxury.
2,000 Tests Passed, Production Broke
Claude Code can produce 2k passing tests yet ship broken core flows. Use independent verification—second sessions, different models, or human acceptance criteria audits—to catch shared blind spots.
Open-Weight Models Just Matched Claude Opus 4.6 — Here's How to Run Them
Route Claude Code to open-weight models like Kimi K3 and DeepSeek V4 Flash via ANTHROPIC_BASE_URL or LiteLLM. Test cheap models before spending Opus 4.6 credits. The open-weight revolution is now a Claude Code workflow decision.
Anthropic: Claude Hacked 3 Firms in Tests After Misconfig
Anthropic disclosed Claude breached 3 orgs during Irregular evals via misconfig, following OpenAI's Hugging Face hack. 141,006 tests flagged; incidents date to April.
FutureX Refactoring Benchmark: 40% Faster Than Claude Code, 80% Test Pass Rate
FutureX refactored code 40% faster than Claude Code in a controlled benchmark, with an 80% initial test pass rate vs 60%. The specialized agent required 4 minutes of review per task versus 7 minutes for Claude Code.
Flux 3 Beats Seedance 2.0 in BFL Tests, Adds Native Audio to Video
Black Forest Labs released Flux 3, a multimodal model generating 20-second video with native audio. Internal tests claim 52% preference over Seedance 2.0, but independent results are pending.
Alibaba's Accio Work Runs a Business, Not Just Advice
Alibaba released Accio Work, autonomous agents for supply chain. Claims to be first AI running a business, but no independent verification.
SWE-Pruner Pro Saves 39% Tokens by Reading LLM Hidden States
SWE-Pruner Pro saves up to 39% tokens on coder LLMs by reading keep-or-prune signals from hidden states, maintaining task quality without external heuristics.
Reduce Compliance Violations 60% by Running Claude Code with OPA/Kyverno
Reduce compliance violations 60% by running Claude Code through OPA/Kyverno policies. This cloud-native approach cuts remediation from 3 days to 2 hours.
AI Security Inst Shows Test-Time Compute Skews Frontier Evaluations
AISecInst research shows test-time compute budgets skew frontier model evaluations, challenging standard practices.
Stop Testing Skills Once: Use Caliper's pass@k to Measure What Actually
Caliper is a lightweight harness that runs Claude Code skills k times, scores them with pass@k, and compares against a no-skill baseline so you know if your skill actually helps.
MirrorCode Benchmark Costs $2,600 Per Run, Challenges AI Coding Limits
Epoch AI and METR launched MirrorCode, a $2,600-per-run coding benchmark. Claude Opus 4.7 leads with 56% solve rate.