# TTFB in AI Search Optimisation: Why We Measure and Optimise (GEO Case Study)

> URL: https://4eck-media.de/en/blog/ttfb-in-ai-search-optimisation-why-we-measure-and-optimise-geo-case-study/  
> Language: en  
> Description: From 1.2 seconds to 184 milliseconds: why the oft-repeated 600-millisecond threshold is a myth, and why we still pay attention to it. A number is circulating in the GEO scene. The…

---

From 1.2 seconds to 184 milliseconds: why the oft-repeated 600-millisecond threshold is a myth, and why we still pay attention to it. A number is circulating in the GEO scene. The claim is that, at a Time to First Byte (TTFB) between 500 and 700 milliseconds, AI bots are increasingly likely to abandon live retrieval. The page falls outside the response budget, the content never makes it into the context, and it is not cited. We examined the evidence, and the result is sobering: there are no real sources for it.

## What can be substantiated, and what cannot

There is no primary source for a 500-to-700-millisecond threshold. Neither OpenAI, Anthropic, Perplexity nor Google publicly document the fetch timeouts of their retrieval bots. What does exist are blog posts that cite one another, while contradicting each other’s figures:

As a [GEO agency](https://4eck-media.de/en/competencies/llm-geo-ai-optimization/), we combine proven SEO fundamentals with optimisation for modern AI search systems.

- One source says that anything below 200 ms is very good, 200 to 500 ms is acceptable, and sustained values above 600 ms demand urgent action.
- Others set 200 ms as the gold standard and claim that citation rates fall significantly above 800 ms.
- Those same texts cite actual per-page timeouts of 1 to 5 seconds, five to twenty-five times the alleged abandonment threshold.

What is striking is where 500 to 700 ms is actually well documented: in conversational AI. There, a median Time to First Token of around 500 ms is considered a practical target for real-time dialogue, derived from turn-taking research; for enterprise voice, the targets are below 400 ms at p50 and below 700 ms at p95. Our assumption is that the number migrated from this context into GEO vocabulary.

Technically, a hard cut-off at 600 ms is implausible anyway. An HTTP client that gave up that early would be absurdly aggressive. Common library defaults are measured in seconds.

## Why TTFB Still Matters: The Race to Be Included in the Context

The fact that the number is wrong does not mean the concern is unfounded. It is just framed incorrectly. In live retrieval, no timeout occurs; there is a **race**. The system retrieves multiple sources in parallel, and the **response budget is finite.** Those who deliver too late are not excluded by an error message; they simply no longer make it into the context from which the answer is built. This is a soft, gradual effect, not a cliff at a particular millisecond value. And TTFB alone is not decisive: so is the time to the last byte of the HTML document, including payload size and compression.

Crawl frequency also matters. Slow respondents are fetched less often, are less current in the index and are therefore less likely to be cited. That has always applied to Googlebot and applies even more to retrieval bots, because their resource budget per domain is calculated more tightly.

The most common causes of failed grounding fetches are not latency at all, but 403 responses from WAF and bot management, JavaScript challenges, 429 rate limiting and client-side rendered markup.

Our position is clear: we are obsessed with PageSpeed anyway. TTFB is a qualifying threshold, not a ranking factor. You have to clear it; then content is what competes. So we use a simple working rule rather than an invented cut-off. Below 200 ms is the target; anything above 700 ms is a genuine problem. And p95 and p99 under load matter, not the median. Retrieval happens at arbitrary times; the worst case decides, not the average.

## The goldmichi.de Case

A current project shows how this plays out in practice. The starting point was a client on an underpowered server, with the corresponding performance figures.

PageSpeed Insights’ Core Web Vitals assessment fails. LCP is 2.6 seconds, in the yellow range; INP and CLS are unremarkable.

    
        
            
                
                    

![Core Web Vitals Vorher](https://4eck-media.de/wp-content/uploads/2026/08/core-web-vitals-vorher.avif "Core Web Vitals Vorher")
                
            
        
    

That week’s TTFB test reveals the cause. The average across European test locations was 1.2 seconds, Grade F and a score of 32 per cent. Even Frankfurt, the geographically closest location, came in at 1.1 seconds. No caching and no offloading beyond the origin.

    
        
            
                
                    

![TTFB vorher](https://4eck-media.de/wp-content/uploads/2026/08/ttfb-vorher-goldmichi.avif "TTFB vorher")
                
            
        
    

The even distribution is noteworthy. There are only 200 milliseconds between the best and worst test location. That is the signature of an origin problem, not a distance problem. A CDN alone would have helped little here because the time was being spent on the server, not on the connection.

After optimisation, the European average is 184 milliseconds. Grade A, a 95 per cent score. Frankfurt is at 128 ms; the slowest test location, in Finland, is at 265 ms, still better than four times what the best location achieved before.

    
        
            
                
                    

![TTFB Nachher, 1. August 2026](https://4eck-media.de/wp-content/uploads/2026/08/ttfb-nachher.avif "TTFB Nachher, 1. August 2026")
                
            
        
    

That is a reduction of around 85 per cent. And because TTFB is the first phase of LCP, the effect carries through directly: the Core Web Vitals assessment now passes, with LCP at 2.2 seconds in the green range.

    
        
            
                
                    

![Core Web Vitals nachher](https://4eck-media.de/wp-content/uploads/2026/08/core-web-vitals-nachher.avif "Core Web Vitals nachher")
                
            
        
    

## Putting the Numbers in Context

Anyone who does the maths will notice that TTFB fell by about one second, while LCP fell by only 0.4 seconds. That is not a contradiction; it is a property of the data source. The PageSpeed Insights figures under “Discover what your real users are experiencing” come from the Chrome User Experience Report and are a rolling 75th-percentile value over 28 days. They therefore still include some measurements from before the optimisation. The effect has not yet fully reached the field data, and LCP will continue to fall in the coming weeks.

As a [web design agency](https://4eck-media.de/en/competencies/ui-ux-web-development-agency-4eck-media/), we translate complex services into understandable content, clear journeys and a robust technical system.

We consider this worth mentioning because case studies often trumpet the maximum delta. The reliable statement is this: the threshold has been cleared, the trend is stable and the remaining improvement will follow automatically.

    
        
            

## Measurably Improve Performance

                            

Further analyses and examples from practice show which technical decisions sustainably improve loading times, Core Web Vitals and perceived quality.

                    
        
                            
                    
                        [![Cleanout: modern brand with one-pager](https://4eck-media.de/wp-content/uploads/2026/02/cleanout-fp-logo-onepager-icons-mapdesign-720x480.avif "Cleanout: modern brand with one-pager")](https://4eck-media.de/en/blog/cleanout-fp-logo-one-pager-with-100-100-seo-96-pagespeed/ "Cleanout F&P: Logo & One-Pager with 100/100 SEO & 96 PageSpeed")
                        
                            [Cleanout F&P: Logo & One-Pager with 100/100 SEO & 96 PageSpeed](https://4eck-media.de/en/blog/cleanout-fp-logo-one-pager-with-100-100-seo-96-pagespeed/)
                                                            
                                    Matthias Petri
                                
                            

                                                    
                    
                

                            
                    
                        [![AVIF: the best file format for your image SEO](https://4eck-media.de/wp-content/uploads/2025/10/squoosh-kompressionsvergleich-webp-avif-bilder-seo-2024-720x480.avif "AVIF: the best file format for your image SEO")](https://4eck-media.de/en/blog/avif-the-best-file-format-for-image-seo-pagespeed/ "AVIF: the best file format for image SEO & PageSpeed")
                        
                            [AVIF: the best file format for image SEO & PageSpeed](https://4eck-media.de/en/blog/avif-the-best-file-format-for-image-seo-pagespeed/)
                                                            
                                    Matthias Petri
                                
                            

                                                    
                    
                

                            
                    
                        [![Fixing CLS problems](https://4eck-media.de/wp-content/uploads/2025/10/cls-probleme-fixen-720x480.avif "Fixing CLS problems")](https://4eck-media.de/en/blog/core-web-vitals-not-passed-identify-and-resolve-cls-issues/ "Core Web Vitals not passed: Identify and resolve CLS issues")
                        
                            [Core Web Vitals not passed: Identify and resolve CLS issues](https://4eck-media.de/en/blog/core-web-vitals-not-passed-identify-and-resolve-cls-issues/)
                                                            
                                    Matthias Petri
                                
                            

                                                    
                    
                

                            
                    
                        [![Average Response Time](https://4eck-media.de/wp-content/uploads/2025/10/average-response-time-720x480.avif "Average Response Time")](https://4eck-media.de/en/blog/improve-ttfb-how-we-were-able-to-reduce-our-servers-initial-response-time-by-88-percent-through-route-caching/ "Improve TTFB: How we were able to reduce our server’s initial response time by 88 percent through route caching.")
                        
                            [Improve TTFB: How we were able to reduce our server’s initial response time by 88 percent through route caching.](https://4eck-media.de/en/blog/improve-ttfb-how-we-were-able-to-reduce-our-servers-initial-response-time-by-88-percent-through-route-caching/)
                                                            
                                    Matthias Petri
                                
                            

                                                    
                    
                

                    
    

## What We Specifically Do in TTFB Optimisation

Goldmichi is a special case: live market values for gold, silver and other precious metals are retrieved every few minutes and directly affect prices in the shop. That is why edge caching is not active here. In principle, however, we take the following measures to improve TTFB in AI Search optimisation:

- Server-side full-page caching rather than rendering on every request
- Object caching for recurring database queries
- Reducing blocking backend calls in the request path
- A current PHP or runtime version and an appropriate OPcache configuration
- An upstream CDN with edge caching for the HTML response
- Cleaning up redirect chains, especially www and protocol redirects

TTFB is a fixed part of our technical AI Search audit, on an equal footing with crawler access and rendering. Specifically, we record:

- TTFB across multiple regions, separated by origin and edge, to distinguish origin problems from distance problems.
- p95 and p99 under load, not the median. The worst case decides the race.
- Time to last byte of the HTML document, because retrieval bots need the complete markup and do not execute JavaScript.
- Log-file analysis by user agent, separating retrieval bots such as OAI-SearchBot, ChatGPT-User, PerplexityBot and Claude-Web from training bots, with status code and request duration for each access.
- Access checks before latency checks: robots.txt, WAF rules, bot score and challenges. A fast page that returns a 403 is worthless.

Point 4 is the only way to answer the question raised at the beginning empirically. If there were an abandonment threshold, it would have to appear as a break in the distribution of aborted connections. So far, we have not seen a published data set that shows this, which is why we analyse it ourselves. Here, for example, is an AI bot log from one of our high-traffic client projects. If requests were being aborted because the server responded too slowly, there would have to be errors with status code 499; we do not have those:

    
        
            
                
                    

![AI Bot Logs](https://4eck-media.de/wp-content/uploads/2026/08/ai-bot-logs-4eck.avif "AI Bot Logs")
                
            
        
    

## Conclusion: The TTFB and AI Search (GEO) Threshold Myth

We do not pass on the 500-to-700-millisecond threshold as fact. It is not evidenced anywhere and is technically framed incorrectly. Still, it would be wrong to dismiss it as unimportant. Whenever we identify TTFB values above 500 ms, we optimise in the client’s interest.

In our view, what remains is another argument for TTFB optimisation, just with the right rationale: faster delivery means more frequent crawling, better odds in the parallel race for the answer context and, as this case shows, passing Core Web Vitals as a welcome side effect. The client received both for the same budget, without changing servers.

Anyone who wants to know where their own website stands should not look at the average, but at p95 under load. That is where it is decided whether you are even in the race.
