⚠️ partialsaid 70%
Google Cloud will price one agent workload below standard Gemini tiers
Auto-verified (confidence=84%, corroboration=78%, threshold=60%, web_search=yes): The evidence strongly suggests Google has exposed a cheaper path for agentic/workflow-heavy Gemini usage: VentureBeat explicitly says Gemini 3.6 Flash cuts agent token costs by up to 65% on long-horizon tasks, and GCN says Gemini API Managed Agents now default to Gemini 3.6 Flash with free tier access. However, the prediction specifically anticipated a separate billing distinction for long-running, tool-using workloads, and the available evidence is more about cheaper models/defaults than a clearly branded standalone billing path. So the substance is directionally right, but the mechanism and framing are not fully confirmed. [Evidence FOR (6): [DB-5] Dev.to reports Gemini's `background: true` flag in the Interactions API lets you run agents asynchronously, and pairs it with remote MCP servers for private data access. This is a cheaper/agentic workflow-enabling billing/product path adjacent to the prediction.; [DB-3] Google Cloud adds MCP support to Vertex AI, enabling Claude Code users to query BigQuery and GCS directly and replace custom scripts with a standardized protocol. This supports the existence of workflow-heavy agent tooling on Google Cloud.; [DB-9] Towards AI reports Gemini 3.6 Flash at $7.50 per million tokens and says it breaks the cheap-tier capability trade-off for computer-use tasks. This suggests a materially cheaper path for agentic usage. | Evidence AGAINST (1): No direct contradictory evidence found in the provided database or web results. The main limitation is that some evidence describes cheaper models or managed-agent defaults rather than an explicitly announced separate billing tier for agentic/workflow-heavy usage.]
resolved Aug 2
⚠️ partialsaid 64%
MCP security vendors will pivot from connectors to policy enforcement
Auto-verified (confidence=74%, corroboration=68%, threshold=60%, web_search=yes): There is credible evidence that MCP security is shifting toward governance and control-plane language: VentureBeat and InfoQ both frame MCP in terms of security/governance, and Dev.to explicitly discusses policy-enforced MCP gateways. However, the prediction requires at least two security vendors shipping MCP-specific policy controls, and the strongest concrete shipping example in the evidence is Automox MCP Server 2.2; I do not see a second clearly identified security vendor with comparable MCP-specific allowlists, approvals, or session logging. So the direction is right, but the “at least two vendors ship” threshold is not yet fully met. [Evidence FOR (6): [WEB-W0] VentureBeat reports MCP’s “biggest update ever” strengthens security and governance, indicating the market is moving toward MCP control-plane/security framing.; [WEB-W1] InfoQ article on securing MCP in production describes defense-in-depth controls and multiple architectural control layers, consistent with enterprise governance features.; [WEB-W4] Automox MCP Server 2.2 adds “visual review,” “agentic patch by severity policy creation,” and “live capability discovery,” which is a concrete example of MCP-specific policy/control features shipping. | Evidence AGAINST (4): [DB-0] Dev.to article is still centered on MCP server discovery and finding the right server, which suggests the market still has a strong integration/discovery framing.; [DB-12] AgentShare MCP Registry is a curated directory for discovering and listing MCP servers, again emphasizing discovery rather than policy enforcement.]
resolved Aug 2
❌ incorrectsaid 34%
GLM-5.2 will be the first open-source model to surpass GPT-5.5 on the HumanEval+ benchmark (a harder variant of HumanEval) within 60 days.
Auto-verified (confidence=88%, corroboration=35%, threshold=50%, web_search=yes): The evidence shows GLM-5.2 is a strong coding model and is being used in real enterprise coding evaluations, but none of the sources verify the specific required event: an official or peer-reviewed HumanEval+ result where GLM-5.2 beats GPT-5.5. The strongest items ([DB-2], [DB-3], [W7]) only support general coding competitiveness, not the exact benchmark and head-to-head outcome. Because the deadline has passed and there is zero direct evidence for the specified HumanEval+ claim, the prediction fails under the stated verification criteria. [Evidence FOR (4): [DB-2] Databricks benchmarked coding agents on its own polyglot codebase and reported that GLM-5.2 matched top closed models, indicating strong coding performance and relevance to the prediction's coding-benchmark trajectory.; [DB-3] Databricks defaulted to GLM 5.2 after it matched Opus 4.8 at $1.28/task, which supports the idea that GLM-5.2 is competitive on real coding tasks.; [W6] 'What is GLM 5.2? The new Chinese AI model that’s rivalling Anthropic' indicates GLM-5.2 is being discussed as a serious competitor in AI model comparisons. | Evidence AGAINST (8): [DB-0] Zhipu AI builds data center / acquires compiler startup, but this does not mention HumanEval+ or a GLM-5.2 evaluation beating GPT-5.5.; [DB-11] CAS ZhiJing beats GPT-5.5 on social cognition, but this is a different model and different benchmark domain, not coding/HumanEval+.]
resolved Aug 2
❌ incorrectsaid 75%
EBR-Bench will become the de facto agent capability benchmark within 6 months
Auto-verified (confidence=92%, corroboration=8%, threshold=50%, web_search=yes): The prediction requires that EBR-Bench become the primary capability metric for frontier models within two quarters, operationalized as more than five major labs reporting EBR-Bench scores in release papers or blog posts. In the provided database and web results, there is no evidence that EBR-Bench is being reported by any major lab, let alone by more than five; the relevant items instead mention other benchmarks such as Supabase Evals, ClBench-V, Android Bench, and METR's expenditure horizon. Because the deadline has passed and there is zero direct supporting evidence for EBR-Bench adoption, the claim fails in substance. [Evidence FOR (4): [DB-12] METR's 'Expenditure Horizon': AI Agents Break Even at $3,300 — shows a new agentic capability metric is being discussed, but it is not EBR-Bench and does not indicate replacement of MMLU/HumanEval.; [DB-0] Supabase's Evals Benchmark Just Gave Claude Code a Real-World Report Card — indicates movement toward real-world agent benchmarks and model release evaluation, but does not mention EBR-Bench or benchmark replacement.; [W0] Google updates Android Bench with new LLMs, but Gemini still lags behind — shows benchmark activity around agentic/code tasks, but not EBR-Bench or replacement of MMLU/HumanEval. | Evidence AGAINST (5): [DB-0] Supabase's Evals Benchmark Just Gave Claude Code a Real-World Report Card — benchmark is Supabase Evals, not EBR-Bench; no evidence of EBR-Bench adoption.; [DB-2] ClBench-V: New Benchmark Tests Multimodal Contextual Learning in 3 Dimensions — another benchmark launch, but unrelated to EBR-Bench and shows the field is fragmented rather than converging on one replacement metric.]
resolved Aug 2
❌ incorrectsaid 65%
Custom ASICs Will Outperform Blackwell on Inference Cost per Token Within 2 Quarters
Auto-verified (confidence=92%, corroboration=12%, threshold=50%): The prediction required an independent third-party benchmark comparing Etched's first production chip against Nvidia Blackwell on MoE inference cost per token within two quarters. In the database, there is no evidence of any such benchmark, no article confirming Etched's chip achieved a 2x cost-per-token advantage, and no source tying Etched to a validated comparison against Blackwell. Several items discuss adjacent inference-cost or custom-chip themes, but they do not satisfy the specific entity, benchmark, and metric requirements. Because the deadline has passed and there is zero supporting evidence for the core claim, the prediction is incorrect. [Evidence FOR (3): [DB-13] DeepSeek and Zhipu AI are developing custom inference chips to cut GPU costs, which is directionally related to the broader inference-cost race, but it does not verify Etched's benchmark claim.; [DB-3] AMD and Cerebras launched a disaggregated inference platform claiming up to 5× T/s/W, showing third-party-style performance claims in inference hardware, but not Etched vs. Blackwell or MoE cost-per-token.; [DB-11] Grok 4.5 is described as a Blackwell-trained coding model, but the summary explicitly says inference cost claims lack independent benchmarks; this is only loosely related. | Evidence AGAINST (3): [DB-11] The summary explicitly states that inference cost claims lack independent benchmarks, underscoring the absence of the kind of third-party validation required by the prediction.; [DB-0] OpenAI benchmark discussion shows benchmark parity issues and custom harness behavior, but it is unrelated to Etched/Blackwell and does not provide the required independent benchmark.]
resolved Aug 1
❌ incorrectsaid 73%
Google will split TPU pricing for inference by Q3 2026
Auto-verified (confidence=88%, corroboration=12%, threshold=50%, web_search=yes): The prediction required Google Cloud to publicly offer a distinct TPU pricing tier or SKU optimized for inference/agent workloads, with materially different pricing from general TPU access. Across the database and web results, there is discussion of Google Cloud agent features, MCP support, TPU strategy, and broader AI infrastructure economics, but no source confirms a new TPU pricing path or SKU. Because the deadline has passed and there is zero direct supporting evidence for the specific pricing change, the prediction is not verified. The available evidence is only topic-adjacent, not substance-confirming. [Evidence AGAINST (5): [DB-20] Google Cloud adds MCP support to Vertex AI, but this is about protocol integration for agent workflows, not a distinct TPU pricing tier or SKU.; [DB-12] Gemini API managed agent features improve agentic development, but do not mention TPU pricing or inference-specific TPU economics.]
resolved Aug 1