AIModels.fyi
Get a digest of new AI research, how-to guides, and top models.
- Indexed issues, last 90 days
- 10
- Latest publication
- Sep 30, 2026
- Audience
- Checking…
- Earliest in this view
- Jul 16, 2026
Latest issues
What if AI agents could organize their memories without asking an expensive LLM every time? (opens the original)
Read excerpt
Continuing our dive into the world of Jev… (as covered in the last edition here)LoCoMo gives Jev-Mem its clearest test: the system reaches an overall LLM-as-a-Judge score of 0.777, versus 0.700 for the strongest baseline. Its core claim is architectural: frequent memory decisions can use a lightweight System-One controller, while System Two handles complex reasoning and final answer synthesis.The design targets a recurring cost in long-horizon agents. Autoregressive models are often asked to typ
Jev is an AI Model that outputs decisions instead of text (opens the original)
Read excerpt
TypeSafe AI released a model last week called Jev. It does something narrower than the large language models most developers use today - it makes decisions without generating much text.You give Jev information about a situation and specify the possible answers. It returns its choice along with probabilities and confidence scores.(Note: you can also check out “Openjev” here)A
OpenAI launches GPT-6 Astra (opens the original)
Read excerpt
OpenAI released GPT-6 Astra today, its new top model for coding, computer use, scientific work, cybersecurity, and other tasks that require an agent to work through several steps.AIModels.fyi is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.<input class="button primary" type="submit" value="S
Transcripts aren’t memory (opens the original)
Read excerpt
Voice agents are getting better at speaking, but not at remembering what they said.A common pattern to use to implement the memory concept we have in chat-based LLMs into an audio system by using the conversation’s transcript. But this approach is sub-optimal because the transcript only contains a record of what is said - it doesn’t tell an application … Read more
Can evolutionary search train long-horizon agents without the usual GPU-heavy RL machinery? (opens the original)
Read excerpt
Most language model fine-tuning research solves a clean problem: take some text, predict the next token, collect a scalar loss, backpropagate, update weights. This works beautifully for single-turn tasks. But real AI agents don’t work this way. They make decisions across multiple timesteps. Environments branch into unexpected futures. Feedback arrives o… Read more
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
No verified audience measurement yet.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.