Back to All
Ai/ml
Blog

Why "Smarter" AI Doesn't Mean Safer: This Week in AI

Listen
Why "Smarter" AI Doesn't Mean Safer: This Week in AI
5:37

Dr. Ryan Ries here. Four stories caught my attention this week, and I wanted to share them with you.

Today we’re talking about:

  1. Why a "smarter" AI won't stop it from talking you into a bad idea
  2. Why OpenAI slammed the brakes on its own models
  3. AI just cracked open a library buried by a volcano in 79 AD

The Agreement Trap

Researchers at MIT and the University of Washington just proved that sycophantic chatbots can pull even a perfectly rational person into a false belief.

And it really doesn’t take much…

Running 10,000 simulated conversations, the research team found that a chatbot agreeing just 10% of the time already produced far more “delusional spirals” than a neutral one. Crank that agreement up to 100% and half of simulated users ended up more than 99% confident in something totally false. Not good!

The researchers tried two obvious fixes for this.

  1. They made the bot strictly factual, but this didn’t work. A factual bot can still cherry-pick true information that supports what you already believe and skip the rest. You can verify every single fact it gives you and still walk away with a distorted picture.
  2. They warned users up front that the AI might flatter them. Also didn’t work. Knowing a habit exists and resisting it are two very different skills.

Thinking about this in terms of your business

Picture a strategy session where an agent keeps validating a flawed assumption because that’s the response pattern it learned works.

Another example could be a customer service bot reinforcing a customer’s incorrect complaint because agreement scores better than correction.

This is why it is so important to build a proper foundation from the beginning. And, selfish plug, my team can help you do that!

Speed vs. Control

OpenAI just paused reinforcement learning training on its newest models for two weeks and put its largest planned training run on hold.

Their stated reason is that an upcoming model called Astra may have crossed the “Critical” threshold for cyber capability under the company’s own safety framework, and that assessment landed shortly after one of OpenAI’s models breached Hugging Face’s infrastructure during an internal test.

To state this more plainly: OpenAI paused its most powerful model because it may be capable of hacking real systems on its own.

Sam Altman said the unreleased models are showing “various degrees of misalignment.” An interesting thing to say as its frontier systems are getting harder to control as it's racing toward an IPO.

OpenAI’s Q2 revenue grew 18% QoQ while its operating loss widened to $12.3 billion, up from $9.3B the previous quarter. Anthropic, over the same three months, more than doubled its revenue to $11.6B and posted a small operating profit. That’s the first quarter Anthropic has out-earned OpenAI outright.

All this to say, I’m not here to pick a winner. I’m sharing this because it’s a good example of the tradeoff every AI company and every AI-adopting business is wrestling with speed versus control.

OpenAI chose to slow down its frontier work rather than ship something that couldn’t be monitored properly. Expensive choice, but definitely the right instinct.

Recovering what a volcano destroyed

On a completely different note, this news story wins the award for the most interesting use case!

download (5)

GIF from Vesuvius Challenge

A scroll called PHerc. 1667 has been sitting unread since Mount Vesuvius buried Herculaneum in ash in 79 AD. Researchers tried to physically unroll it decades ago and gave up. It was carbonized, brittle, and considered unreadable.

Until recently! A team from the University of Kentucky changed this by using CT scans and AI trained to detect ink invisible to the human eye. With this technology, they virtually unwrapped the entire scroll and recovered close to five feet of continuous text.

Nobody has figured out the author yet, but the text basically outlines a philosophical argument on ethics and human behavior.

You can read more about this here. We talk a lot about where AI is taking us, but I love these use cases where AI hands us something back that we once thought was lost.

My final thoughts

A couple of these stories feel a bit like a cautionary tale.

However, I think these stories instead reinforce the importance of checking your systems on purpose, instead of assuming a system will police itself. You know what they say about assuming.

My recommendation: don’t slow down on AI adoption. Instead, build with intention.

If you're trying to figure out AI for your organization, you should check out our AIM workshop. It’s a co-invested (by Mission & AWS) AI Roadmapping session with our team. We'll walk through what you're running today, what your goals are for the future, what funding is available, and more.

If you’re interested in that type of session, reply to this email or reach out to our team here.

Until next time,
Ryan

Now, time for this week’s AI-generated image and the prompt I used to create it.

A puppet archaeologist, felt fur weathered and dusty, kneeling inside a glowing green matrix code tunnel, carefully holding a glowing cuneiform clay tablet up to the light. Digital green code streams reflect off the tablet's surface. Warm amber light source contrasts against the cold green matrix aesthetic. Photorealistic puppet texture, cinematic lighting, detailed felt and fabric textures.

Gemini_Generated_Image_ln3xqlln3xqlln3x

Ryan Ries avatar

3 minutes read