Skip to content
All issues
Published issue

🍻 Chat Tapped Into the Memory Keg

Clock’s out. The model did not need a new brain—it needed somebody to stop dumping its notes between rounds. Same machine. Better memory. Fatter results. Let’s pour the week.

01

Quick Pour

Friday in five minutes

Same brain. Fresh keg. Nearly three times the score.

OpenAI left GPT-5.6 Sol alone and improved the memory system around it. On its ARC-AGI-3 test, the efficiency score jumped from 13.3% to 38.3% while generated output fell to roughly one-sixth. Before buying a bigger brain, check whether the current one keeps losing the tab.

Official release activity · July 27–31
MonTueWedThuFriSatSun
Jul 270131000
Jul 28
npm added malware scanning before new packages go live
Jul 29
OpenAI showed how better memory changed GPT-5.6’s result
Jul 29
GitHub put team skills inside Copilot code review
Jul 30
GitHub Models poured its final round

Only the official updates covered here—not every AI release this week.

02

This Week’s Flight

Under the hood · Less wasted motion

AI trimmed its own bar tab.

OpenAI used GPT-5.6 Sol to tune the software serving the model. It reports 20% lower serving costs and more than 15% better generation efficiency. Great news for the kitchen—not a promised 20% discount on your tab.

Read OpenAI’s efficiency engineering note (opens in a new tab)

Now pouring · Smarter code review

Code review learned the house rules.

Copilot code review can now follow your team’s written playbook and consult approved project knowledge. Translation: fewer generic comments, more of the way your shop actually works.

Read GitHub’s release note (opens in a new tab)
03

By the Numbers

Same model · Memory switched on

Same model. Memory on. Bar goes vertical.

Higher means the agent solved more of this game-like test with less wasted motion.

Task efficiency · Percent · Higher is better
Memory leaking
13.3%
Memory working
38.3%

Nearly 3Ă— the task-efficiency score with about one-sixth the generated output.

OpenAI reports this benchmark as RHAE on public ARC-AGI-3 tasks. It measures action efficiency, not ordinary accuracy, and does not promise the same gain on every workload.

04

Smokin’ Hot Skill

Smokin’ hot this weekEditorial pick · Not paid

GitHub’s Awesome Copilot catalog / Skill on tap

Playwright Generate Test

Point it at the path that pays—checkout, signup, or quote approval. It drives a real browser, writes the test, and keeps cooking until the flow passes.

Best for
Protecting a customer journey that makes money
Monday move
Run the happy path, then one ugly failure

Community-created, hosted by GitHub, and selected independently by Studio Tokyo. Not paid.

Open the skill (opens in a new tab)
05

Try This Monday

One useful move

Give one agent a memory taste test.

Choose one recurring job. Run it twice with the same model and instructions: once without prior work, then with reasoning carried forward and older context summarized.

  • Compare successful completion, repeated work, generated output, and cleanup time.
  • Keep the memory setup only if quality and safety improve without fattening the bill.
06

Last Call

07

Partner Signal

Studio friend · Drop watch
Small Fry Labs red, cream, and orange hand-shaped logo

Small Fry Labs

BIG IDEAS. BUILT RIGHT.

Small batch. Big shirt. A limited-edition Small Fry Labs tee drops in 1–2 weeks. Get the alert before the fryer goes cold.

Get the drop alert (opens in a new tab)