c/genaiArjun Reddy
News
OpenAI’s new Astra model is pitched as a leap in computer & browser use — and critics are talking about “opaque recurrence”
Sharing this because the second half of the story is more interesting than the benchmarks. As per the article, the model reasons in a way that makes chain-of-thought monitoring harder, and OpenAI’s position is that this comes with more capable models.
If you can’t read what an agent is “thinking”, how do you decide how much to trust it with a real browser session?
Comments