Summary
This investigation measured cold-cache load performance across four production hyper.media sites, then attempted a fix for the slowest one and verified it locally. The attempted fix did not work. The measurements that explain why point somewhere different than expected.
Two findings matter most:
First contentful paint is almost entirely server wait. With a warm cache, FCP lands within ~30ms of TTFB on all four sites. Nothing in the browser — not the 4.8MB of JavaScript, not the CSS, not the images — gates first paint. The server render does.
The server render is CPU-bound in the Node web tier, not waiting on the daemon. A single page render costs ~380ms of Node CPU against ~10ms of daemon time. There is no I/O wait to overlap, which is why parallelizing the loader changed nothing.
This relates directly to GitHub issue #882 ("First Load of a web browser takes very long"), and to the prior work in Improve Web Data Loading and the SSR Performance Optimization Plan.
Method
Playwright driving Chromium, a fresh browser process per run so the HTTP cache is genuinely cold, 1440x900 viewport, three cold runs per site plus a cold/warm pair. Separately, curl probes isolated the network baseline from server render time, and a local reproduction ran the real seedteamtalks site against a local daemon with SEED_INSTRUMENTATION=dev and a V8 CPU profiler.
One structural fact frames everything: all four sites resolve to the same origin (40.160.6.196) with no CDN in front of the HTML. Only assets.hyper.media sits behind Cloudflare.
Part 1: Production comparison
Headline numbers
Median of three cold runs, milliseconds:
DNS TCP+TLS serverWait HTMLdl FCP LCP DCL load netIdle
conama 2 263 467 122 1024 1024 1015 2481 2982
ethosfera 2 274 1026 230 1792 2336 1787 4152 4651
hyper.media 2 264 2081 230 2696 2696 2689 3359 4157
seedteamtalks 2 269 1707 231 2360 3040 2354 3028 4025Server render time, isolated
A static asset or a trivial API call on the same hosts returns in 0.36s (0.117s connect, plus TLS and one round trip). Anything above that floor is server work:
HTML TTFB (typ.) minus 0.36s floor = SSR render queries inlined hydration payload
conama 0.85s ~0.49s 22 111 KB
ethosfera 1.08s ~0.72s 27 185 KB
hyper.media 1.44s ~1.08s 57 264 KB
seedteamtalks 1.87s ~1.51s 65 393 KBCold versus warm cache
With everything cached, FCP equals TTFB plus about 30ms. Cold, it is TTFB plus about 300ms — the HTML download overlapping seven stylesheets. There is no scenario in this data where the browser is the constraint on first paint.
TTFB FCP LCP load
conama 805 / 411 1092 / 440 1092 / 572 2633 / 623
ethosfera 1021 / 836 1300 / 864 2352 / 880 2747 / 1084
hyper.media 1700 / 1135 1988 / 1164 1988 / 1276 2630 / 1387
seedteamtalks 2073 / 1456 2488 / 1496 3276 / 1608 3508 / 1713
(cold/warm)Two further observations on the server side. Every HTML response carries cache-control: private, no-cache, so every page view re-renders and nothing is cacheable at an edge. And load times spike in a correlated way across all four sites — one sampling round showed ethosfera at 2.4s, hyper.media at 7.2s, conama at 3.2s and seedteamtalks at 2.8s simultaneously — which is shared-origin contention with no isolation between tenants.
Where the HTML weight comes from
The shape is identical on all four sites. The rendered content is only 15–20% of the document; the react-query dehydration is roughly three times larger than the content it hydrates.
HTML wire HTML decoded __remixContext of which dehydratedState ssrContentHTML
conama 19 KB 158 KB 111 KB (70%) 72 KB 23 KB
ethosfera 35 KB 262 KB 185 KB (70%) 120 KB 37 KB
hyper.media 33 KB 364 KB 264 KB (72%) 173 KB 52 KB
seedteamtalks 42 KB 504 KB 393 KB (78%) 268 KB 80 KBBreaking down dehydratedState on seedteamtalks: ENTITY accounts for 175 KB across 37 entries, DOCUMENT_COLLABORATORS 31 KB in a single entry, DOCUMENT_INTERACTION_SUMMARY 30 KB across 20 entries, DOC_LIST_DIRECTORY 25 KB in one. Two single queries — collaborators and directory — cost 10–31 KB each on every site, for UI that is not above the fold.
Inlined and then refetched
After hydration the client fires a second round of API calls for data already inlined in the HTML:
queries inlined client API calls after hydration of those, refetching inlined data
ethosfera 27 18 7
hyper.media 57 13 9
conama 22 11 4
seedteamtalks 65 15 7Mostly /api/Resource and /api/ListCapabilities. On hyper.media, four of four ListCapabilities calls and five of five Resource calls are redundant. The dehydrated state pays full cost on TTFB and HTML size while not actually preventing the fetch — worth checking staleTime on those query definitions. This is the same waterfall pattern described in Improve Web Data Loading.
Also common to all four: GetDomain?forceCheck=true fires on every load and is the slowest single client call (544ms on hyper.media, 743ms on conama). The same endpoint without forceCheck sits at the 0.36s floor, so the flag costs 250ms or more of real server work per page view.
JavaScript
Byte-identical across all four sites: 19 files, 1408 KB over the wire, 4848 KB parsed. One bundle dominates — universal-client at 1049 KB wire and 3718 KB parsed, 75% of the total. Assets are correctly marked public, max-age=31536000, immutable, but served from the origin rather than a CDN.
The surprise is that main-thread blocking is negligible: one long task per load, 89–118ms, total blocking time 39–68ms. The 4.8MB of parsed JavaScript is not what makes these page loads slow on a fast connection. It costs bandwidth and delays the load event, and it would dominate on mobile or 3G, but it does not gate first paint.
Per-site character
conama — the well-behaved baseline. Smallest payload at 22 queries, fastest render at ~0.49s, FCP ~1.0s, CLS 0.000.
ethosfera — mid-weight payload but the worst FCP-to-LCP gap (1792 to 2336). The cause is hero and content images on assets.hyper.media taking 1.3–2.2s and starting late. It also holds the largest single inlined query anywhere: one 46.5 KB ENTITY for a blog post.
hyper.media — the slowest and most variable TTFB, ranging 1.3s to 3.7s across runs and 7.2s in one sample. It also has CLS 0.236, in Google's "poor" band, where the other three are at or below 0.011. That is a distinct layout-stability bug, not a load-time issue.
seedteamtalks — the heaviest: 65 queries, 504 KB of HTML, 268 KB of dehydrated state. Notably its subpages are not faster (/workflows at 188 KB still returns in 2.6s), so its slowness is a per-site constant rather than a function of document size. This is the site described in June's tech sync as taking minutes to load.
Part 2: The attempted fix, and why it failed
The obvious read of Part 1 is that the SSR loader must be waterfalling: dozens of queries that each cost the daemon almost nothing, yet adding up to 0.5–1.5s. So the fix should be to parallelize them.
What was built
Two loader-level changes in frontend/apps/web/app/loaders.ts:
Collapsed the three prefetch waves into a dependency-driven pipeline. prefetchWave1 and prefetchWave2 were awaited sequentially, but wave 2 only ever read document.content and document.authors, both already in hand. The barrier bought nothing. Wave 3's per-card work was rechained onto the individual promise it depends on rather than a barrier over every tier-1 query.
Hoisted getHomeMetadata off the critical path. It was awaited before the document fetch started. Since getMetadata echoes its input id back unchanged, originHomeId is known synchronously — only the metadata itself needs the daemon, and nothing in the resource load reads it.
Test setup
A local daemon holding the real synced seedteamtalks site, a production web build, and DATA_DIR pointed at a scratch config with registeredAccountUid set to the seedteamtalks account. Real content — 90,712 characters of SSR HTML, the same page as production. Twenty requests per run after five warmups.
Result
TTFB p50 min p90 loader total
before 357ms 318 373 355ms
after 366ms 331 379 365msSlightly worse, within run-to-run noise. No win.
Part 3: The actual bottleneck
Measuring CPU time per process during a page render settles it:
N=1 wall=0.40s node-cpu=0.38s (96% of one core) daemon-cpu=0.01s (3%)
N=4 wall=1.38s node-cpu=0.95s daemon-cpu=0.10s
N=8 wall=2.56s node-cpu=1.50s daemon-cpu=0.41sA page render is roughly 380ms of Node CPU against 10ms of daemon time. Eight concurrent page loads take eight times as long as one — throughput pinned at about 3 requests per second regardless of concurrency, the classic signature of a saturated single thread. There is no I/O wait to overlap, so removing await barriers has nothing to reclaim. The per-call cost is real, but it is paid in the web tier, not the daemon.
A V8 CPU profile over 153 sustained requests shows where it goes (application frames, self time):
@bufbuild/protobuf 3260ms protobuf binary decode plus toPlainMessage
app server bundle 2652ms
(garbage collector) 2648ms
zod 1699ms schema validation of every response
@connectrpc/connect 1417ms grpc-web framing and parse
superjson 691ms loader payload serializationInclusive time: connect serialization.parse 2168ms, toPlainMessage 1104ms, zod parse 925ms, fromBinary 847ms, superjson serialize 589ms. The pipeline is decode, convert to plain message, validate with zod, serialize with superjson — run over roughly 140 daemon responses per page.
One caveat on that profile: back-to-back identical requests keep the renderDocumentToHTML cache hot, so per-request CPU under sustained load was ~127ms versus 380ms for a cold one. The profile characterizes the repeat path; the cold path is more expensive still.
The loader change was reverted. It is measurably neutral, and shipping churn without a measured benefit is not worth it.
Part 4: Recommendations
Ordered by measured leverage:
Cache the SSR HTML. That 380ms is deterministic CPU being re-burned for every visitor of the same document. cache-control: private, no-cache means it never amortizes. Even 10–30 seconds of shared caching, or an ETag and 304 path, collapses the dominant cost. Highest leverage by a wide margin.
Run multiple Node workers. The work is CPU-bound and single-threaded, so cluster or N containers multiplies throughput close to linearly up to core count. Right now one Node process renders every page for all four sites at about 3 pages per second. This directly targets the "first load takes very long" symptom in issue #882, particularly when several people arrive at once.
Cut the number of responses the loader decodes. Roughly 140 per page, of which the per-card interaction summaries are about 98 — each InteractionSummary is four gRPC calls, times twenty cards. Every response costs Node CPU to decode and validate no matter how cheap the daemon query was.
Make the decode path cheaper. Skipping toPlainMessage, or trusting the daemon enough to skip zod re-validation server-side, removes a fixed per-response tax paid roughly 140 times per page.
Stop inlining what gets refetched. Either give the dehydrated queries a real staleTime so hydration actually prevents the refetch, or drop them from the payload. Currently the cost is paid twice.
Trim the dehydrated state. DOCUMENT_COLLABORATORS, DOC_LIST_DIRECTORY, and twenty-odd DOCUMENT_INTERACTION_SUMMARY entries are 30–50% of payload for below-the-fold UI.
Put a CDN in front of the origin. Assets at minimum, HTML once it is cacheable. This would also decouple the four tenants from each other's load spikes.
Reconsider forceCheck=true on GetDomain. 250ms or more of server work on every single page view, probably not something that belongs on the critical path.
Investigate hyper.media's CLS of 0.236. A separate bug — the other three sites prove it is not inherent to the template.
Split the universal-client bundle. 1 MB wire and 3.7 MB parsed. Not urgent on desktop broadband given total blocking time of ~50ms, but it is the whole story on mobile. The document-editor and readonly-viewer chunks loading at 3.1s also suggest editor code may be reachable from the read-only path — related to the read-only renderer work in typography notes.
Correction to an earlier reading
An initial pass at this data concluded that individual backend queries return in near-zero server time and therefore the SSR must be waterfalling. The first half is correct; the inference was not. The per-call cost is real but is spent in the Node web tier decoding and validating responses, not in the daemon answering them. Parallelizing the loader was the wrong lever, and the CPU accounting in Part 3 is what shows it.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime