I want faster generation when it shortens the path to finished work. That is why Gemma 4 drafters interest me. In May 2026, Google announced speculative decoding with generation up to three times faster without losing quality, according to its evaluation. That is a concrete improvement in one part of execution.
Looking back at the announcement in October 2026, I remember what I reported in July: an agent solved something in minutes after it had been taking the whole day. I was excited because the work finally moved. The episode identifies neither a model nor the accelerated stage, so it cannot serve as a Gemma test. It explains the result that makes me impatient for the next improvement.
In September 2026, I requested an investigation into speeding up development without losing quality. I want more experiences like that July case, with less waiting between intention and delivery. A technique promising to preserve quality deserves attention because I can examine it without starting by accepting a worse answer.
I still need to locate the delay. Text generation occupies one portion; tool research and deciding how to approach a problem occupy others. The experience of speed combines them all. If most waiting happens outside generation, that stage can improve considerably without producing an equivalent change in total task time.
For a brief request, waiting for the answer may dominate the experience. For an investigation with several tool calls, the rest of the journey deserves more attention. These are examples for a comparison I still want to perform. The same technical gain can help both and represent very different shares of the wait I observe. The vendor's number directs my attention.
What I can do with the released time changes too. An earlier answer lets me correct direction earlier if I am available. When I leave a long task running while working on something else, I pay more attention to the whole job finishing. The improvement has value through the use I can make of it, including when I am away from the screen.
My next evaluation will repeat a type of request with a defined starting point and usable result. I will record where waiting occurs and check quality before attributing improvement to generation. The July case remains my reference for excitement: work got unstuck. That is the effect I will look for when trying a speed technique.