🤖 AI

GPT-6 Reuses Earlier Prompt Text to Make Long AI Jobs Faster and Cheaper

3 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
prompt caching

A system that reuses a repeated part of a prompt instead of processing it from scratch.

cache hit rate

The share of requests that successfully reuse cached input.

persistent agent

An AI system that keeps working through a long, multi-step task.

What changed

OpenAI announced an improved prompt caching system for GPT-6 on September 22, 2026. OpenAI’s announcement says GPT-6 can reuse eligible shared prompt prefixes within a 30-minute window. Cached input tokens can receive discounts of up to 90 percent.

GPT-6 is designed for persistent agents. These agents may work for hours on coding, research, documents, or presentations. Their applications send many API requests during one task. Those requests often repeat the same instructions, tool definitions, and earlier context.

Why the background matters

A prompt is more than the user’s newest question. It can include rules, reference material, tool descriptions, and conversation history. If an application sends the same long beginning repeatedly, the model may process that material repeatedly as well.

Prompt caching lets an application reuse a shared beginning instead. It is not permanent memory. It is a short-term way to reuse eligible context. A changing ending can be added after a stable beginning. That makes caching especially relevant for multi-step agents.

Why this matters

Long-running agents need many model requests. Each request can add waiting time and input-processing costs. Reusing stable context can reduce both. It can also make it easier for developers to run background tasks or continue several branches of one conversation.

The larger change is operational. Developers no longer have to guess why caching stopped working. OpenAI added tools for measuring cache use and investigating misses.

What OpenAI confirmed

A new Prompt Caching Dashboard shows how much application input came from the cache. Developers can track hit rates over time and compare cached and uncached tokens.

A diagnostics tool compares a request with a recent response. It can identify changes to the model, tools, settings, or input. It also estimates how many reusable tokens were affected.

Developers can place explicit cache breakpoints. These mark which prompt prefix should remain reusable. GPT-6 also allows developers to change reasoning effort between responses without breaking the earlier cache. They can append a configuration_update while keeping the request-level setting unchanged.

OpenAI also recommends keeping tool definitions, schemas, and ordering stable. Developers can add new instructions later in the context. They can prewarm shared instructions and reference material before a user asks a question.

OpenAI cites several early results. It says GitHub Copilot reduced the share of prompt tokens requiring fresh processing by more than 50 percent against an earlier baseline. It also presents partner examples. One company reported costs falling 20 percent. Manus reported cache hit rates rising from about 85 percent to above 90 percent. Wordsmith reported a rise from 83 percent to 91 percent, with inference costs falling 36 percent.

What remains unknown

These numbers come from OpenAI and participating companies. They are not a broad independent benchmark. Other applications may see smaller gains. Results will depend on prompt stability, tool changes, request patterns, and task length.

The 90 percent figure applies to cached input tokens. It does not mean every request becomes 90 percent cheaper. A cache miss can still happen when important context changes. The 30-minute reuse window also means this system is different from permanent memory.

What to watch next

Developers will need to measure cache hit rate, response speed, and real bills in their own applications. The key question is whether the new controls make long-running agents reliably cheaper and faster in everyday use. OpenAI’s prompt caching guide gives the technical details, but real workloads will show how much of the promised benefit appears outside the early examples.

🤖 AI

GPT-6 Can Reuse Earlier Explanations

📰 Full story: GPT-6 Reuses Earlier Prompt Text to Make Long AI Jobs Faster and Cheaper

A new caching system may make long AI jobs faster and less expensive.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
prompt cache

A saved part of a prompt that the model can use again.

cache miss

A case where saved prompt text cannot be reused.

cache breakpoint

A marker that shows where reusable prompt text ends.

💡 The gist

  • GPT-6 can reuse parts of earlier prompts.
  • This can speed up long, multi-step AI jobs.
  • Developers get tools to find cache misses.

OpenAI, the company behind GPT-6, improved prompt caching. Its announcement explains the change.

A prompt is the text sent to an AI model. It can include instructions, tool descriptions, and earlier conversation. Long AI jobs often send the same information many times.

A prompt cache saves a repeated beginning. That beginning is called a shared prefix. GPT-6 can reuse it when a later request matches. This means the model does not need to process every repeated word again.

OpenAI says cached input tokens can receive discounts of up to 90 percent. The discount applies to cached input. It does not mean the whole job always costs 90 percent less.

Prompt caching is not permanent memory. The saved text is useful for a limited time. OpenAI describes reuse within a 30-minute window. Changed tools, settings, or input can prevent reuse. That failure is called a cache miss.

Developers now get a dashboard. It shows how much input came from the cache. A diagnostics tool can explain a miss. It can point to changed tools, settings, models, or text.

Developers can also add a cache breakpoint. This marks where reusable text ends. GPT-6 can change its reasoning effort between responses while keeping earlier context reusable. Developers can also prepare common context before the user asks a question.

OpenAI reports better results from several partners. Some reported higher cache hit rates and lower costs. These examples do not guarantee the same results for every app. Each app sends different prompts and uses different tools.

The next step is testing. Developers should compare speed, cache hit rate, and actual bills. The improvement reduces repeated work. It does not remove every cost from a long AI task.

The official caching guide explains the available controls in more detail.

🤖 AI

GPT-6 Can Reuse the Same Note

📰 Full story: GPT-6 Reuses Earlier Prompt Text to Make Long AI Jobs Faster and Cheaper

GPT-6 may work faster when it receives the same long note again.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
GPT-6

A computer system that reads text and gives answers.

prompt

A note that tells an AI what to do.

cache

A nearby saved copy of repeated information.

OpenAI, the company that makes GPT-6, added a helper. OpenAI explains it here.

GPT-6 is a computer that reads and answers. It often receives long notes. Those notes are called prompts. A prompt tells the computer what to do.

GPT-6 can keep a repeated beginning nearby. That nearby copy is called a cache. A cache is like a shelf for a note. It is not forever memory.

GPT-6 may reuse the note within 30 minutes. Then it does less repeated reading. The answer may arrive faster. Reading costs for cached input may drop by up to 90 percent.

OpenAI made a screen for developers. The screen shows whether the cache helped. It can also show when the cache was not used. That is called a cache miss.

Developers can mark which part to reuse. Changing the note can stop reuse. The helper saves repeated work for long computer jobs.

Sources