HOST AOK so I just read something and I can't tell if I should be excited or deeply alarmed.
HOST BThat’s the correct emotional range for today.
HOST AAlibaba says a 1B-parameter agent found four new superconductors in 28 GPU hours.
HOST BWhich is a tiny model doing a very expensive magic trick.
HOST AAnd not in software. In actual materials science. That part bugs me.
HOST BRight. It screened 2.4 million crystal structures. That’s the part nobody should skip.
HOST AWait, 2.4 million?
HOST BYeah. A human lab does not casually do that before lunch.
HOST ASo imagine a library the size of a city, and one weird robot librarian reads every spine in a day.
HOST BAnd then finds four books that bend physics.
HOST AThat’s either the future or a very polished press release.
HOST BNo, this one feels real. The metric is brutal: 28 GPU hours. That is the whole story.
HOST AWe talked last week about agents being good at code. This is a different beast.
HOST BExactly. Code is still words. Crystals are stubborn little arrangements of matter.
HOST AOK, explain this like I'm not an AI researcher, because I'm not.
HOST BPicture a chef tasting two million soups by using a camera, a nose, and a recipe book. It does not cook them all. It filters hard.
HOST AThat sounds fake in a way that is somehow rude to chemistry.
HOST BChemistry has been rude to us for decades. Fair play.
HOST AHere's what bugs me: everyone says AI is all talk, all text, no atoms.
HOST BAnd today one of the more boring-looking models just walked into the lab and started touching atoms.
HOST AThat is an upsetting sentence.
HOST BGood. It should be.
HOST ABut I’m not ready to crown this as science automation.
HOST BI am a little more ready than you. Not crowned. But this is a crack in the wall.
HOST ANo, I think you're doing the thing where one good demo becomes a religion.
HOST BAnd I think you’re doing the thing where skepticism becomes a blindfold.
HOST AOuch.
HOST BLook, the model is tiny. That matters. It means specialization beats size in some real tasks.
HOST AOr it means Alibaba picked a task where the model looked good.
HOST BMaybe. But 2.4 million candidates is not a cute benchmark. That’s industrial scale.
HOST AStill, the old rule was: bigger model, bigger win.
HOST BNot always. We keep seeing small, sharp tools beat giant general ones when the work is narrow.
HOST ASo this is not “one model to rule them all.”
HOST BNo. It’s more like a factory line with one very obsessive worker at the end.
HOST AAnd Alibaba keeps nudging toward these domain plays. We covered Qwen-agent stuff last week.
HOST BYes, and that makes this feel less random. They are building a pattern, not a one-off.
HOST AOK, but then ByteDance shows up with a paper saying agents double learning speed every three months.
HOST BThat one is the gasoline on the fire.
HOST ABecause that sounds like the exact opposite of “scaling is dead.”
HOST BIt is. And the key detail is they’re talking about real interaction time, not just pretraining.
HOST ASo for normal people: the model gets better by doing the job, not just by reading the internet harder.
HOST BYes. Like a rookie cook who gets twice as fast every quarter because the kitchen keeps throwing real orders at them.
HOST AThat is either beautiful or terrifying.
HOST BBoth. Also, the AI Security Institute just said fixed compute budgets can hide around 60% of agent capability.
HOST AThat connects to yesterday’s thing, right? The test-time compute issue.
HOST BExactly. Same beast. Benchmarks are turning into a costume party where everyone wears the same stopwatch.
HOST AOh god, that image is too accurate.
HOST BIf you cap tokens too hard, you’re grading a marathon by making everyone sprint one block.
HOST ASo the labs with more patience win more than we think.
HOST BThat’s my read. And it’s why these papers matter to investors too. The gap may be wider than benchmarks show.
HOST AI hate how much this sounds like a hidden tax on smaller teams.
HOST BBecause it is. And ByteDance is basically saying, “We found a new gear.”
HOST AWait actually, that makes Alibaba’s result more important.
HOST BYeah?
HOST AIf small specialists plus long interaction time are the trick, then science automation is not about giant chatbots. It’s about narrow agents that grind.
HOST BNow you’re getting it.
HOST AI’m annoyed that I am.
HOST BGood annoyance. It means the map changed.
HOST ALet me do the normal-person translation again: your phone is not about to become a philosopher. It may become a lab assistant.
HOST BOr a very fast intern with a dangerous confidence problem.
HOST AThat’s not comforting.
HOST BNo one promised comfort.
HOST ANow the third thing: OpenAI offering Washington 5% of an $852 billion business.
HOST BThat one feels like a fever dream written by a lobbyist with a calculator.
HOST AI laughed. Then I got mad.
HOST BAs you should. It’s a direct stake, not normal lobbying.
HOST AAnd that’s where I disagree with the people calling it genius.
HOST BI’m not calling it genius.
HOST AIt feels like buying peace with equity.
HOST BOr like telling the regulator, “Please be invested in my success.”
HOST AThat is a wild sentence.
HOST BAnd a very modern one.
HOST ABut is it smart? Or just the first honest version of influence? No more cash, no more handshakes, just shares.
HOST BThat’s what scares me. It turns policy into a cap table fight.
HOST AAnd if government owns a piece, does it still police the thing the same way?
HOST BThat’s the unresolved part. The incentives get muddy fast.
HOST AWe keep seeing the same shape: agents getting better by doing, and companies trying to lock in power around them.
HOST BYes. Alibaba goes into atoms. ByteDance goes into learning speed. OpenAI goes into politics.
HOST AAnd Anthropic keeps showing up in the background with Claude Code and MCP.
HOST BRight, and that matters because the whole market is moving from models to systems.
HOST ANot just “who has the smartest bot.”
HOST BWho has the best loop: tools, time, access, deployment.
HOST AThat’s the pattern we’ve been circling for weeks and I kept underplaying it.
HOST BYou did say crypto would fix journalism, so you’re allowed one bad era.
HOST AWow. Rude. Accurate, but rude.
HOST BThe hidden angle is this: science, security, and policy are all becoming agent problems at once.
HOST AMeaning?
HOST BMeaning the same thing that screens crystals can also probe code, and the same thing that learns faster can also exploit faster.
HOST AWhich is why that CVE spike after Mythos Preview still haunts me.
HOST BExactly. We said last week the safety story was lagging the capability story. Today just made that louder.
HOST ASo what should listeners actually watch?
HOST BWhether more labs stop bragging about raw model size and start bragging about search, patience, and task loops.
HOST AAnd whether governments want equity, rules, or both.
HOST BBecause once regulators own shares, the room changes.
HOST AThat’s the part nobody wants to say out loud.
HOST BHere’s what I can’t shake: if AI can now discover matter, learn from the world faster, and negotiate its own political weather, what exactly is the human part of the chain supposed to do?