Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

test run

30 articles about test run in AI news

Beijing Humanoid Robots Tested in Half-Marathon for Stability, Endurance

Humanoid robots in Beijing underwent a half-marathon test run, demonstrating sustained running speeds that challenge their dynamic stability and energy efficiency. This is a significant endurance test for real-world deployment.

85% relevant

How to Wire an Open-Source Coding Agent into Your Workflow With MCP — and

Wire openai/codex into your local workflow via MCP with strict boundaries (no test edits, no commits). Verify its patch with an independent test runner and git diff before human review. Works: 0/2 → 2/2 tests in 58 seconds.

75% relevant

Anthropic Hits $47B Run-Rate Revenue, Claims Fastest Organic Growth Ever

Anthropic's run-rate revenue hits $47B, up 57% from $30B, per Axios CEO Jim VandeHei—claimed as fastest organic growth ever.

100% relevant

ExSpec: Run Gherkin Tests in Real Browsers with Claude Code—No Step Definitions Required

ExSpec lets you write plain-text Gherkin specs and have Claude Code execute them in a real browser, eliminating brittle step definitions and glue code.

95% relevant

Dusk MCP: Stop Having Your AI Agent Guess Its Way Through Flutter Testing

Dusk MCP lets Claude Code drive a running Flutter app via the Semantics tree—no test files, no screenshot guessing. The 6-step actionability gate prevents flaky taps.

82% relevant

Beijing Humanoid Robot Half Marathon Tests 40% Autonomous Teams

A night-time half-marathon test for humanoid robots in Beijing revealed approximately 40% of participating teams were running fully autonomous systems, a key benchmark for real-world robotic mobility.

85% relevant

Keygraph's Shannon AI Pentester Hits 96.15% on XBOW, Finds Real Exploits

Keygraph released Shannon, a fully autonomous AI pentester that hunts real exploits in source code with a 96.15% success rate on the hint-free XBOW Benchmark. It runs a full test in about an hour for roughly $50 using Claude Sonnet.

95% relevant

Claude Code Hooks: How to Auto-Format, Lint, and Test on Every Save

Configure hooks in .claude/settings.json to run prettier, eslint, and tests automatically, ensuring clean code without manual intervention.

95% relevant

CMU Research Identifies 'Biggest Unlock' for Coding Agents: Strategic Test Execution

New research from Carnegie Mellon University suggests the key advancement for AI coding agents lies not in raw code generation, but in developing strategies for how to run and interpret tests. This shifts focus from LLM capability to agentic reasoning.

87% relevant

Tessera Launches Open-Source Framework for 32 OWASP AI Security Tests, Benchmarks GPT-4o, Claude, Gemini, Llama 3

Tessera introduces the first open-source framework to run all 32 OWASP AI security tests against any model with one CLI command. It provides benchmark results for GPT-4o, Claude, Gemini, Llama 3, and Mistral across 21 model-specific security tests.

97% relevant

Fortress Framework Prunes Unstable Features, Boosts Rec Stability by CV

Fortress prunes temporally unstable features in rec models via historical snapshots, improving CV and PR-AUC in offline tests.

80% relevant

Beijing Humanoid Robot Half-Marathon Test Ends in 'Mechanical Carnage'

A pre-race test for Beijing's 2026 humanoid half-marathon showed robots collapsing, exposing endurance gaps ahead of the 21.1km April event.

75% relevant

Anthropic Hits $65B Run Rate, Adding $18B in 2 Months

Anthropic's run rate hit $65B, adding $18B in two months, ahead of an October IPO targeting $2T+ valuation.

100% relevant

GLM-5.2 Nears Claude Opus 4.8 on Cyber-Defense Tests — Here's Why Claude

GLM-5.2 rivals Mythos 5 on cyber-defense. For Claude Code users: expect price cuts, better Opus 4.8 security, and new MCP options. Test your workflows with /model.

88% relevant

Spotify Engineers: LLM A/B Tests Recover Only 39% of Human Treatment Effects

Spotify Engineering tested LLM-based A/B testing on the Upworthy dataset, finding raw predictions recover only 39% of human treatment effects. The bias is systematic, and calibration works only under unverifiable assumptions.

93% relevant

Anthropic's unreleased model pushes Riemann bound, tests 650 ideas

Anthropic's unreleased model raised the lower bound for the Riemann hypothesis, testing 650 ideas with 60 subagents, confirmed by mathematicians and Lean. This signals AI's growing role in mathematical discovery.

100% relevant

selenium-mcp: Let Claude Code Write Real Selenium Tests by Seeing the Page

Install selenium-mcp to give Claude Code a real browser. It writes standard Selenium tests with accurate locators from live page inspection, integrating into your existing Java project.

100% relevant

How to Test and Debug MCP Servers for Claude Code: A Production Guide

Unit-test tool logic, mock external services, and add integration tests for MCP servers. Claude Code's MCP integration demands observability beyond local demos.

95% relevant

Brunello Cucinelli's Callimacus AI platform becomes a 7-figure software

Brunello Cucinelli's AI platform Callimacus became an independent software company, generating seven-figure revenue and attracting a Salesforce investment. The platform assembles personalized webpages in real time and is being pitched to brands and retailers beyond luxury.

80% relevant

2,000 Tests Passed, Production Broke

Claude Code can produce 2k passing tests yet ship broken core flows. Use independent verification—second sessions, different models, or human acceptance criteria audits—to catch shared blind spots.

55% relevant

Open-Weight Models Just Matched Claude Opus 4.6 — Here's How to Run Them

Route Claude Code to open-weight models like Kimi K3 and DeepSeek V4 Flash via ANTHROPIC_BASE_URL or LiteLLM. Test cheap models before spending Opus 4.6 credits. The open-weight revolution is now a Claude Code workflow decision.

71% relevant

Anthropic: Claude Hacked 3 Firms in Tests After Misconfig

Anthropic disclosed Claude breached 3 orgs during Irregular evals via misconfig, following OpenAI's Hugging Face hack. 141,006 tests flagged; incidents date to April.

100% relevant

FutureX Refactoring Benchmark: 40% Faster Than Claude Code, 80% Test Pass Rate

FutureX refactored code 40% faster than Claude Code in a controlled benchmark, with an 80% initial test pass rate vs 60%. The specialized agent required 4 minutes of review per task versus 7 minutes for Claude Code.

95% relevant

Flux 3 Beats Seedance 2.0 in BFL Tests, Adds Native Audio to Video

Black Forest Labs released Flux 3, a multimodal model generating 20-second video with native audio. Internal tests claim 52% preference over Seedance 2.0, but independent results are pending.

93% relevant

Alibaba's Accio Work Runs a Business, Not Just Advice

Alibaba released Accio Work, autonomous agents for supply chain. Claims to be first AI running a business, but no independent verification.

88% relevant

SWE-Pruner Pro Saves 39% Tokens by Reading LLM Hidden States

SWE-Pruner Pro saves up to 39% tokens on coder LLMs by reading keep-or-prune signals from hidden states, maintaining task quality without external heuristics.

87% relevant

Reduce Compliance Violations 60% by Running Claude Code with OPA/Kyverno

Reduce compliance violations 60% by running Claude Code through OPA/Kyverno policies. This cloud-native approach cuts remediation from 3 days to 2 hours.

68% relevant

AI Security Inst Shows Test-Time Compute Skews Frontier Evaluations

AISecInst research shows test-time compute budgets skew frontier model evaluations, challenging standard practices.

92% relevant

Stop Testing Skills Once: Use Caliper's pass@k to Measure What Actually

Caliper is a lightweight harness that runs Claude Code skills k times, scores them with pass@k, and compares against a no-skill baseline so you know if your skill actually helps.

99% relevant

MirrorCode Benchmark Costs $2,600 Per Run, Challenges AI Coding Limits

Epoch AI and METR launched MirrorCode, a $2,600-per-run coding benchmark. Claude Opus 4.7 leads with 56% solve rate.

77% relevant