Case Study

How NerdWallet Eliminated 10-Second GraphQL Query Planning Delays

SoundCloud logo with a white cloud and waveform inside a purple hexagon pattern background.

10s → <1s

Query planning time on the slowest operations

Order of magnitude faster

Faster on previously slow operations after warmup

3-phase migration

Phased migration designed to protect production traffic

Industry
Music Streaming
Company Size
700+

Company Overview

NerdWallet is a personal finance platform that helps people make smarter decisions about their money by comparing and evaluating products like credit cards, banking, mortgages, investing, and insurance. It offers tools, content, and a mobile app to help consumers understand options, choose financial products, and track their finances with greater confidence.

The Challenge

A Gateway Architecture Under Pressure

Before switching to WunderGraph, SoundCloud relied on another federated gateway. It covered the basics, but going deeper into its features would have tied them more closely to that system than they wanted.

Upgrading was blocked.

‍
Legacy code tied the team to an old Federation version. New subgraphs followed current Apollo patterns, but the gateway couldn't validate or plan those shapes. Teams building subgraphs found themselves constrained by a router that couldn't keep up.

‍Query planning started blocking traffic.

‍
Complex queries took more than 10 seconds to plan. While those queries waited, other requests backed up behind them. The slowdown spread across the graph and began affecting users.

To compensate, the team wrote their own internal prewarming logic. It helped until a new slow query appeared. If that query wasn't added to the list, the performance of the entire graph could suffer.

    Building it ourselves was something we talked about, but the amount of effort required to build and maintain it long term just wasn't worth it

    Tim Caplis, Principal Engineer

    The Solution

    A Phased Migration to WunderGraph Cosmo

    When NerdWallet evaluated alternatives, they focused on three things: Federation support, performance, and cost, both in infrastructure and team time. The migration also had to be safe. 

    They chose WunderGraph Cosmo and ran a three-phase rollout:

    ‍Match the gateway. Ensure Cosmo Router supported every feature the Apollo Gateway provided before touching traffic.
    ‍
    ‍Sync schema registries. Publish schema updates to both registries simultaneously so the old gateway and Cosmo could stay aligned during testing.
    ‍
    ‍Shift traffic gradually. Move production traffic from Apollo Gateway to Cosmo Router in small increments, validating at each step.

    When roadblocks appeared - missing Federation features, behavior gaps tied to newer Apollo spec versions - the WunderGraph team patched them or built the missing support directly.

    The WunderGraph team jumped in to patch issues and build features that unblocked our migration.

    Subgraph teams retained ownership of their services and features throughout the migration. Internal QA ran regression tests at each stage. The phased approach allowed them to migrate while keeping user-facing behavior stable.

    See Cosmo in action

    Learn more about Cosmo

    The Results

    Lower Costs, Better Performance, Greater Flexibility

    Cosmo produced immediate, measurable improvements:

    • Cache warming eliminated repeated 10+ second planning delays and brought planning latency down dramatically after warmup.

    • Removed fragile internal prewarming scripts that required manual updates every time a slow query was added

    • High confidence in Supergraph updates and horizontal scaling, with router version upgrades still handled carefullygs: $265,000

    • Improved visibility into slow queries and bottlenecks via Cosmo Studio and OTEL tracing

    • Reduced operational overhead — no more emergency prewarm list maintenance, fewer brittle workarounds around the old gateway

    Our infrastructure costs went from $14,000 with our previous provider down to $9,750 with Cosmo.

    Tim Caplis, Principal Engineer

    The Conclusion

    Cache Warmer: The Performance Breakthrough

    The biggest improvement came from Cosmo's Cache Warmer.

    Previously, NerdWallet's internal script tried to prewarm the worst offenders. It worked, but it required manual updates.

    Cosmo Cache Warmer replaced this entirely. It automatically identifies and prewarms the most expensive query plans before serving traffic. Subsequent requests skip the planning phase completely.

    "Caching query planning was a game-changer in terms of performance for us."

    For some of their slowest production operations, cache warming eliminated repeated 10+ second planning delays and dramatically improved performance after warmup. Similar cold-start planning behavior was later documented publicly in WunderGraph’s Super Bowl scaling analysis, where cache warm-up reduced planning spikes from 8-15 seconds to below one second.

    Improved Observability Across the Graph

    ‍
    As the platform grew, the team needed better visibility into what was happening inside the graph.Cosmo Studio provided traces, analytics, and schema visibility in one place. The team could see slow requests, identify complex execution paths, and quickly spot where subgraphs were struggling, turning guesswork into targeted fixes.OTEL tagging confirmed that performance stayed on par or improved through the migration. Cache-hit metrics helped tune the router after rollout.

    See what WunderGraph can do for you!

    Speak to an expert

    Talk to Our Technical Experts

    Tell us about your GraphQL Federation setup and what you're trying to accomplish. We'll get back to you with practical next steps and answer any technical questions you have.

    By clicking "Get in touch", I acknowledge I have read and understand the Privacy Policy.
    Oops! Something went wrong while submitting the form.

    Trusted by platform teams. Loved by developers.

    Frequently Asked Question

    What led SoundCloud to migrate away from their previous GraphQL gateway?

    SoundCloud’s third-party gateway worked initially but introduced high infrastructure costs and risk of lock-in due to enterprise-only features. They needed more flexibility and control over their GraphQL infrastructure without restrictive licensing.

    How did WunderGraph Cosmo help SoundCloud reduce infrastructure costs?

    Cosmo reduced CPU usage by 86%, from 600 vCPUs to 80, saving an estimated $265,000 annually. Their overall infrastructure spend dropped from ~$14,000 to ~$9,750 per month even after accounting for other infrastructure components.

    What performance improvements did SoundCloud see after switching to Cosmo?

    One high-traffic query related to reactions dropped in latency from 171ms to 94ms (P95). Routing became more efficient, reducing bottlenecks and improving responsiveness across views like track waveforms and comment counts.

    How did SoundCloud approach the Cosmo migration process?

    They ran Cosmo alongside their existing gateway to validate schema checks and measure performance. Once confident, they phased out the old gateway and fully transitioned to Cosmo Router.

    Why did SoundCloud prefer an open-source federation router?

    Cosmo’s open-source model gave SoundCloud full control without enterprise restrictions. They could inspect the codebase, avoid vendor lock-in, and maintain flexibility in how they manage their GraphQL architecture.

    Blue circular badge with text 'AICPA SOC' and a URL for soc4so on aicpa.org.
    SOC 2 Type II
    Certified

    Ready to scale without the chaos?

    Speak with an expert to discover how Federation becomes simple and seamless.