HOST AOK so I just read something and I can't tell if I should be excited or annoyed.
HOST BThat's my favorite kind of AI news.
HOST ADeepSeek shipped V4 Flash 0731 at $0.14 per million input tokens, and it scored 50.
HOST BAnd GPT-5.6 Luna is still ahead, but not by much.
HOST ARight. This is not a tiny gap anymore.
HOST BIt's like a beat-up hatchback keeping up with a luxury sedan.
HOST AExcept the hatchback costs way less and keeps getting better.
HOST BAnd we said last week DeepSeek was building for scale in Inner Mongolia.
HOST AYes, and now the model itself is the story.
HOST BWait, the bigger shock is Frontis-MA1.
HOST AOh god, yes.
HOST BA 35B meta-evolution agent scoring 71.21 on MLE-Bench Lite, above GPT-5.5 and Codex.
HOST AIf that holds up, the whole 'just make it bigger' religion gets a bad day.
HOST BCan I be real for a second? That would be a punch in the face to the scaling crowd.
HOST AAnd maybe deserved.
HOST BMaybe. Or maybe it's a benchmark doing benchmark cosplay.
HOST ANo, that's a garbage take if you stop there.
HOST BI'm listening.
HOST AThe pattern is not 'small model wins once.' The pattern is training method starting to matter more than raw size.
HOST BOr it's one very specific test where the model learned the trick.
HOST ASure, but that's still the point. Tricks compound.
HOST BNope. Not enough. We have to be careful here.
HOST AFine. Here's what bugs me: we keep treating compute like the only ladder.
HOST BAnd the ladder is now getting weird rungs.
HOST AExactly. DeepSeek's ten-point jump without a disclosed size jump suggests efficiency, not brute force.
HOST BThat part scares me more than the price.
HOST AExplain this to me like I'm not an AI researcher, because I'm not.
HOST BOK so imagine two kitchens. One buys a bigger stove. The other learns to waste less food and plates better.
HOST AAnd suddenly the second kitchen is beating the first on taste and bill.
HOST BExactly. That's DeepSeek right now.
HOST ABut Frontis is stranger, because it hints at a model improving itself.
HOST BRecursive self-improvement sounds like a sci-fi poster until it wins a benchmark.
HOST AOr until it doesn't and everyone quietly deletes the press release.
HOST BFair. The claim is unverified, so I'm not marrying it.
HOST AGood, because I was about to throw the whole thing in the trash.
HOST BDon't. Just don't worship it.
HOST AI won't. But if a 35B model can do this, then model size is becoming one input, not the whole game.
HOST BThat's the ugly truth for the big labs.
HOST AAnd for the people buying API access by the million tokens.
HOST BWhich brings us back to DeepSeek: cheap intelligence is getting less fake.
HOST AAnd that matters right now because pricing used to be the moat.
HOST BNow it's more like a speed bump.
HOST AWait, actually, maybe the moat is trust and distribution, not raw model quality.
HOST BMaybe. But if quality keeps flattening across price tiers, the trust moat gets crowded fast.
HOST AWe should say this for normal people: cheaper models mean more apps can use AI without melting their budget.
HOST BAnd also more junk can be built faster. Same coin, different face.
HOST AThat's the part nobody puts in the launch post.
HOST BNo, they put it in the funding deck.
HOST AHa. OK, but the hidden thread today is not 'cheap.' It's 'structured.'
HOST BKeep going.
HOST ADeepSeek is squeezing more intelligence out of fewer dollars. Frontis is squeezing more performance out of training process. Meta is squeezing knowledge out of papers into atomic claims.
HOST BWait, that last one is the quietest and maybe the smartest.
HOST AMeta's AskChem turned 147,000 chemistry papers into 2.4 million cited claims.
HOST BThat is not a search box. That's a claim factory.
HOST AExactly. Instead of handing you chunks of text and hoping, it hands you little facts with a DOI attached.
HOST BFor people who don't dream in Python: this is like going from a pile of books to a catalog of answers.
HOST AAnd that connects straight back to the model story.
HOST BBecause now the battle is not just who can talk best.
HOST AIt's who can organize reality better.
HOST BThat sentence should be illegal.
HOST AI know. But it's true.
HOST BAnd AskChem is a knowledge graph move, not a chatbot move.
HOST AYes. It's less 'ask me anything' and more 'here are the pieces, already labeled.'
HOST BWhich is why researchers will love it and regular users will never know it exists.
HOST AUntil every product quietly starts doing this behind the curtain.
HOST BAnd that is the same story as DeepSeek and Frontis.
HOST ARight — AI is moving from big monoliths to systems that do one part of the job absurdly well.
HOST BWe said a month ago the frontier would stop looking like one giant model race.
HOST AAnd now it looks like a bunch of specialist knives.
HOST BSharp, cheap, and slightly unsettling.
HOST AWhich is honestly a very on-brand sentence for 2026.
HOST BI hate that you're right.
HOST ASo my takeaway is simple: don't stare only at the leaderboard. Watch what happens when models get cheaper, narrower, and better at one thing.
HOST BMy takeaway is harsher: the old belief that scale alone wins is now a theory under pressure.
HOST AAnd the pressure is coming from both sides.
HOST BFrom cheaper models on one side, and from smarter training on the other.
HOST APlus the search layer, which everyone ignores until it changes how science gets done.
HOST BRemember when we joked that the future would be one giant autocomplete machine?
HOST AYeah.
HOST BTurns out it's also a filing clerk.
HOST AThat's the most depressing compliment I've ever heard.
HOST BI mean it lovingly.
HOST AWhat I can't stop thinking about is this: if intelligence is getting cheaper and more structured, what exactly are we paying for next week?