0:00 / 3:32
Chapters
Sources
DAILY ROUNDUP
OpenAI's Jalapeño Chip Beats Nvidia On Inference Speed And Power
calendar_today Date:
schedule Duration: 3:32
visibility 5 Views
OpenAI's Jalapeño chip delivers up to 3.6x lower latency than comparison systems, Claude unifies memory across chat and Cowork, Ollama simplifies Claude Desktop gateway switching, and humanoid robots race a chaotic obstacle course in Beijing.
- 01. Jalapeño was tested across GPT, DeepSeek and Kimi models, delivering up to 4.1x higher performance on interactive workloads without trading throughput for latency
- 02. Claude's unified memory works both ways between chat and Cowork, on by default across Free, Pro and Max on web, desktop and mobile
- 03. Ollama v0.33's one-toggle gateway lets Claude Desktop use any local or cloud model inside Ollama, then switch back with no filesystem edits
- 04. The World Humanoid Robot Games obstacle course saw one robot catch fire and several fold at the waist after the finish line, with a person still on the controls
Today's AI news: OpenAI shared the first real test results for Jalapeño, its custom inference chip, delivering up to 1.9x more AI work per watt and up to 3.6x lower latency than comparison systems. Claude now has one shared memory across chat and Claude Cowork. Ollama v0.33 makes connecting Claude Desktop to Ollama as a third-party gateway a single toggle. And humanoid robots raced a chaotic 100-metre obstacle course at the World Humanoid Robot Games in Beijing.
Chapters:
0:00 Today's AI News
0:25 OpenAI Jalapeño Chip
1:18 Claude Unified Memory
2:01 Ollama Claude Desktop Gateway
2:47 Robot Obstacle Course
Blend Roundup 2026-08-26
https://x.com/OpenAI/status/2092300846675505602
https://x.com/claudeai/status/2092299704864284888
https://x.com/ollama/status/2092453536634380763
https://x.com/HumanoidsHQ/status/2092142977158185314
[quick] OpenAI's Jalapeño chip beat Nvidia GB300 on inference latency and power, Claude now shares one memory across chat and Cowork, Ollama made switching Claude Desktop to local models trivial, and humanoid robots raced an obstacle course in Beijing. [excited] Here's today's AI news.
[quick] OpenAI has shared the first real test results for Jalapeño, its custom inference chip, and the numbers are striking - up to 1.9 times more AI work per watt and up to 3.6 times lower latency than comparison systems.
For genuinely interactive workloads, where response speed matters most, OpenAI measured up to 4.1 times higher performance - all in one architecture, rather than trading throughput for latency the way most hardware has to.
It was tested across GPT, DeepSeek and Kimi models, so this isn't just cherry-picked for OpenAI's own performance.
OpenAI plans to start deploying Jalapeño in its own infrastructure by the end of this year, alongside its existing Nvidia hardware, with a second and third generation already in development behind it.
[quick] Claude now has one shared memory across chat and Claude Cowork, and you're the one who decides what stays in it.
Hand a task to Cowork and it starts from what Claude already knows from your chats - the project you talked through, your manager's preferences, or a client from last quarter - instead of you re-explaining everything from scratch.
It works both ways too, so anything Cowork picks up on a task feeds straight back into your regular chats, and Claude builds this up live as you talk, rather than only summarising once a conversation ends.
It's on by default across Free, Pro and Max, on web, desktop and mobile, and everything it remembers is listed out for you to review, edit or delete.
Ollama version 0.33 has landed, and it makes connecting Claude Desktop to Ollama as a third-party gateway a single toggle instead of a manual setup job.
Once it's on, you can pick from any model inside Ollama, whether that's running locally on your own machine or hosted on Ollama's cloud, right from within Claude Desktop itself.
Turn it off again and you're straight back to your normal first-party Anthropic setup, no filesystem edits or config files involved either way.
It's a genuinely simple fix for something that used to take real manual work, and it puts local and cloud open models properly on equal footing with Claude inside the same app.
Robots raced a hundred-metre obstacle course at the World Humanoid Robot Games in Beijing, and it's every bit as chaotic as it sounds.
The course packs in L-turns, a crawl section, a vault, a spiral staircase and much more, all designed to test genuinely dynamic, varied movement rather than just walking in a straight line.
It wasn't clean - one robot collided with an obstacle and caught fire, and several others couldn't stop after the finish line and folded at the waist, so this is still very much a work in progress.
A person is still on the controls for the obstacle course rather than the robots running it fully autonomously, which is worth remembering next time the highlight clips make it look effortless.
Meta Data
Model: