A pretty cache hit rate can hide lazy work distribution. If every agent receives material it does not need, paying less for repeated reads improves the bill while preserving the mess. I want to see both before celebrating a percentage.
On August 20, 2026, OpenAI launched a prompt caching dashboard showing hit rates and consumption broken down into tokens read from cache, written, and left uncached. That visibility interests me because it helps locate repeated context. The dashboard supplies evidence for a decision; deciding whether the material belongs in a request still takes judgment.
In June, I had objected to passing the context of every project between agents. Around the same time, I suspected there were too many agents and suggested reducing the number. I was looking at the organization of the execution and questioning whether that distribution made sense. A consumption chart could help locate the problem without deciding responsibilities for me.
Stable repository instructions may be necessary across several calls. The complete history of a discussion about another assignment has a much weaker justification. Packaging everything together because it is easy leaves the agent to separate what matters. Later, somebody still has to establish why the bill grew. I prefer examining one concrete context block to issuing a blanket instruction to shorten every prompt.
I also need comparable executions. Changing the instructions or the documents received can alter consumption without any change to the model. If the record omits that, the investigation starts by assigning the cause to the wrong place. A cheap day and an expensive day tell me little without knowing what work was happening.
Isolation remains useful even when it costs some repetition. Several agents may need common rules; decisions from unrelated assignments can confuse who should act on what. I will pay for information needed to keep that responsibility clear. A caching discount does not persuade me to spread context merely to increase reuse.
I will use the dashboard to investigate one apparently irrelevant context block. After establishing its purpose, I want to compare executions with and without it, checking cost alongside the deliverable. If the bill falls and the result worsens, both belong in the record. I want the cache hit rate to help choose what to send next. Once the percentage becomes the objective, it gets easy to optimize reading while forgetting the work that justified the call.