The Cloud Is Getting Faster. Does That Make It Ready for Latency-Sensitive Market Data?
The Cloud Is Getting Faster. Does That Make It Ready for Latency-Sensitive Market Data?

For years, the trade-off in market data infrastructure was easy to describe. Colocation was fast but expensive and rigid. Cloud was flexible and operationally convenient but too far from the action for anything latency-sensitive. That line is getting harder to draw.
CME Group is in the process of migrating Globex markets to a private Google Cloud region built around Google’s Ultra Low Latency (ULL) infrastructure. Once the migration is complete, firms will be able to choose between traditional self-managed colocation and Google’s specialized cloud infrastructure-as-a-service, with both options offering equal network latency to the exchange. That is a meaningful shift. It means the cloud is no longer automatically the slower option at the point where it matters most: the connection to the matching engine.
It is tempting to read that as “cloud has caught up, so the colocation era is over.” That framing misses the more useful question. Equalizing network latency at the exchange edge does not automatically make an entire market data pipeline low latency. It changes where the starting line is. What happens after that starting line, across capture, normalization, distribution, storage, and analytics, is still an architecture decision every firm has to make on its own.
What Actually Benefits From the Cloud
Cloud infrastructure is genuinely good at a specific set of jobs. Elastic compute for backtesting and research. Storage that scales without a procurement cycle. Distribution to teams and applications that are not latency-sensitive at all, like compliance, reporting, or historical analysis. Cloud also removes a lot of operational overhead: patching, hardware refresh cycles, physical capacity planning. None of that changes because the exchange happens to be nearby in network terms.
Where cloud infrastructure struggles is anywhere that consistent, sub-millisecond behavior is the requirement rather than the goal. Shared infrastructure, virtualization layers, and general-purpose networking were not built to guarantee the same round trip every single time. ULL offerings like Google’s are narrowing that gap for the network hop itself. They are not, on their own, solving determinism further down the stack.
Where Physical Proximity Still Matters
For market data ingestion specifically, proximity matters most at the point of first contact with the feed: the feed handler that receives raw market data off the wire. This is where jitter and packet loss have the biggest downstream impact, because every microsecond of delay or inconsistency here compounds through everything built on top of it. That does not make the feed handler the only proximity-sensitive piece of a firm’s infrastructure. Execution and strategy components can carry equally or more stringent proximity requirements depending on the use case. This section is specifically about the ingestion layer.
It matters less the further a workload gets from that first contact point. A risk engine consuming a normalized, already-timestamped feed does not need to sit next to the exchange. Neither does a research notebook querying historical order book data. Proximity is not a blanket requirement for the whole stack. It is a requirement for specific components, and the further you move down the pipeline, the less it tends to apply.
Network Latency Is Not the Whole Story
This is the distinction that gets lost in most “cloud versus colocation” conversations: network latency and application latency are two different problems, and equalizing one does not equalize the other.
- Network latency is the time it takes a packet to travel between the exchange and your infrastructure. This is what CME and Google are addressing with equal-latency access.
- Application latency is everything that happens once the data arrives: parsing, decoding, normalization, sequencing, timestamping, and handing the message off to whatever consumes it next.
A feed handler sitting one hop from the matching engine can still erase some or all of the latency advantage that proximity was supposed to provide, if the software itself is inefficient, if it is competing for CPU cycles with other workloads, or if it is built on a stack that was never designed for raw, high-frequency delivery. Moving infrastructure closer to the exchange buys back network time. It does not automatically buy back processing time, and processing time remains one of the more controllable variables in the latency budget.
The Case for Hybrid Architecture
Once network and application latency are treated as separate problems, a hybrid architecture stops looking like a compromise and starts looking like the logical design. Capture and initial distribution sit as close to the source as the use case demands, on infrastructure built specifically to handle raw feed delivery with minimal overhead. Analytics, storage, and less time-sensitive distribution sit wherever is most cost-effective and operationally convenient, which today is very often the cloud.
This is closer to how firms serious about market data have approached the problem for years, even before ULL cloud options existed. Raw feed handling stays lean, dedicated, and physically close to the source. Everything built on top of that feed, from analytics platforms to internal tools to AI-driven research, can live anywhere that makes operational sense, because it is consuming clean, already-normalized data rather than fighting for proximity it does not actually need.
Predictable Beats Lowest Average
Infrastructure teams often optimize for the lowest average latency, which is a reasonable instinct but an incomplete one. For most latency-sensitive market data use cases, predictability matters more than a marginally lower average. A feed handler that delivers consistently at 40 microseconds is more useful than one that averages 25 microseconds but occasionally spikes to 400. Tail latency and jitter are what break downstream systems, whether that is a trading strategy, a risk check, or a data pipeline that assumes messages arrive in order and on time.
Raw, low-level delivery protocols still have an important role even as network infrastructure improves. Direct UDP or TCP delivery can reduce unnecessary abstraction between the source and the consuming application, giving engineering teams greater control over how data is received, processed, and distributed. That control also makes it easier to place latency-sensitive feed handling exactly where the application’s latency budget requires, rather than being locked into a single vendor’s colocation footprint.
None of that is automatic. UDP or TCP delivery does not create predictability on its own. Buffering strategy, threading model, kernel and NIC configuration, and how the receiving application parses and sequences messages still determine whether that control actually translates into consistent, low-latency performance.
What to Measure Before Deciding Where a Feed Handler Lives
- Actual latency budget by use case. Not every application needs the same latency profile, and treating them all as equally sensitive leads to over-engineering in some places and under-engineering in others.
- Jitter and tail latency, not just average latency, measured under realistic load rather than idle conditions.
- Behaviour under peak message rates. Measure latency, packet loss, queue depth, and processing stability during market-open and volatility-driven bursts rather than relying on steady-state benchmarks.
- Where processing overhead is actually being introduced. Profile the feed handling and normalization layer separately from the network hop, so the two are not conflated when something is slow.
- Failover and redundancy requirements, since a design that is fast under normal conditions but fragile during a failover event has simply moved the risk rather than removed it.
- The true cost of proximity, including colocation fees, cross-connects, and specialized hardware, weighed against what is actually gained for each specific workload.
None of this is a case against the cloud, and none of it is a case against colocation. It is a case for measuring before deciding, component by component, rather than picking one infrastructure model and applying it uniformly across a pipeline that has very different requirements at different stages.
Closing Thought
The CME and Google Cloud migration is a useful signal for the industry, not because it settles the cloud versus colocation debate, but because it shows how thin that debate has become. When network latency is no longer the differentiator, the conversation has to move to what actually is: how data is captured, how it is processed, how consistently it behaves under load, and where each piece of that chain genuinely needs to live.
The question is no longer simply cloud or co-location. The better question is: where should each component of the market data pipeline live?
That is the design question infrastructure teams should be asking now, regardless of which cloud or colocation option they eventually choose.

