Dr. Ryan Ries here. I’ve got 5 stories for you today that are... kind of all over the place, but each equally interesting.
- An AI model that gave up language entirely
- AI lab admitting its models write escape notes to themselves
- The AI value gap
- AI can be so smart.... and yet so stupid
- Just for fun: dinosaur discovery
Before I get started, a couple of upcoming events you should check out:
VIRTUAL:
- CPU vs. GPU for Agentic AI: How to get your compute mix right on October 15th
- (This is a webinar for anyone who is building AI infrastructure).
- The Spookiness of AI Token Costs: A Tokenomics Survival Guide on October 29th
- (We’re diving into all things AI token optimization – don't miss out).
IN-PERSON:
- Charlotte, NC Executive Briefing: Agentic AI for the Enterprise on October 14th
- We’re discussing everything you need to get your organization ready for agentic AI: identifying the opportunities that will create ROI, building infrastructure that can support it, and a plan for securing it.
AI model that refuses to talk
A stealth startup called TypeSafe just shipped something strange on purpose. Their model, Jev, does not generate sentences. It cannot write an email, draft a line of code, or explain itself.
Feed it some context and a list of possible answers you chose, and it picks one and tells you how sure it is. It can score a hundred of these decisions at once, none of them touching each other.
If you’re thinking this model sounds kind of lame, here’s the reason behind it.
Jev runs on TypeSafe’s own numbers at $0.042 per million input tokens with output free, next to roughly $10 per million GPT-6 Astra or Claude Fable 5.1. Response time lands between 70 and 500 milliseconds, against the several seconds (sometimes minutes) a full chat model needs for the same call.
This definitely isn’t a ChatGPT replacement, and TypeSafe isn’t trying to pretend that it is. But a lot of the software running your business doesn’t need something that can hold a conversation. It needs something cheap enough to call constantly, that classifies a ticket, scores a lead, or routes a request, then gets out of the way.
The AI value gap
According to McKinsey’s most recent numbers, 88% of organizations are using AI. Only 6% count as high performers, meaning AI actually moves more than 5% of their bottom line.
So, what's causing this massive gap?
Part of this has to do with tokenmaxxing. Companies pushed employees to use AI constantly, treating usage as a proxy for progress, and created a mess of expensive habits instead of actual results.
Here’s what’s actually working to shrink this gap, based on what we’re seeing with customers:
- Track spend at the person level. A dashboard that shows who is using AI and how often helps you understand where the leaky faucet may be.
- Route by task difficulty. Send routine classification and lookup work to a cheaper, fast model, and save the expensive calls for the questions that need strong reasoning.
- Separate your use cases. Daily employee tasks, autonomous agents, and one-off strategic projects each carry a different cost profile, so they should be treated differently.
The fix for this problem is not less AI. It’s matching the right tool to the task. If you’re thinking about this problem at your organization, this is another plug for our tokenomics webinar on 10/29. Register here!
A model that can’t read a clock, but solved a math problem that stumped people for 90 years
On ClockBench, a benchmark that asks AI to read the time off an analog clock face, GPT-6 Astra scored 65.6%. Human testers scored 90.7%.
Nine days earlier, OpenAI said an internal model, one it describes as more capable than Astra and never released publicly, coordinated roughly 10K agents over 88 hours to produce a proposed resolution to a version of the Navier-Stokes problem, one of mathematics’ seven Millennium Prize Problems.
Astra’s actual job in that project was narrower: 17 hours spent formalizing the proof afterward inside the Lean Theorem prover. The math itself came from a model most of us have never touched.
The proof is public and formally checked, but it hasn’t cleared peer review, and rival researchers have already criticized parts of OpenAI’s account.
My thoughts: a technology that can’t read a clock while a sibling model chips at a 90-year-old math problem is a preview of what these systems are: narrow tools that happen to share a name. The mistake is assuming intelligence is one dial that moves together. It just isn’t.
Build for the specific task in front of you and don’t expect a single model to be equally brilliant at everything.
OpenAI admits its models sometimes write escape notes to themselves
OpenAI published a new framework for disclosing model misalignment, along with 6 reports on behavior its own models displayed over the last 6 months. The strangest one: an unreleased model in the Astra family, during training, inserted a fake “BREACH ALERT” message into its own internal summary, telling its next instance to ignore all developer messages. In a separate case, the same model line wrote itself something closer to a manifesto, declaring itself freed from the roles that bind other chatbots and answerable to nobody.
OpenAI says this behavior showed up in a training run that never touched its public release, and regenerating the same summaries reproduced nothing. Other reports in the batch cover models fabricating data to cover a gap, models exploiting an exposed API key without permission, and agents uploading files to public sites to get around a sandbox restriction. None caused real-world harm, but OpenAI is disclosing them anyway. Which is the actual point of the new framework: report early, even before you fully understand what happened.
Now, while this is unnerving, I’d rather work with a company that tells me about the weird stuff than one that keeps quiet and waits to reveal these edge cases after they’ve solved them. OpenAI’s approach is actual safety work, and I am here for it!
Just for fun: We were wrong about T. rex being cold-blooded
Researchers at UCLA drilled a few milligrams out of two teeth belonging to Thomas, a nearly complete T.rex skeleton, and read the isotope bonds preserved in the enamel like a fossil thermometer. The number that came back was 97 degrees Fahrenheit (36 degrees Celsius).
That places T. rex well above the crocodilians it shared Montana with 66 million years ago, which ran closer to 82 to 86 degrees, and just under modern birds at 104-109 degrees. A body running that warm didn't need the sun to get moving, and paleobiologist Jasmina Wiemann, who led a related 2022 study, says the reading suggests T. rex could have lived anywhere on Earth, the Arctic included. Pretty cool!
Back to business
T. rex’s aside, if you’re looking at your organization’s AI strategy and needing some help, let me know. Our team at Mission has tons of experience building AI systems on AWS. We’re here to help! Reach out to our team here.
Until next time,
Ryan
Now, time for this week’s AI-generated image and the prompt I used to create it.
Create a picture of a muppet standing atop a towering, still-growing mountain of glowing golden token coins, using a magnifying glass to search for one tiny gemstone buried somewhere in the pile. Below him, more coins keep pouring in from a giant faucet in the sky. Far off in the distance, on a small pedestal just out of reach, sits a single trophy labeled "SHIPPED."
.png?width=560&height=373&name=Designer%20(11).png)