Scale-Up vs Scale-Out: How UALink and Ethernet Divide the AI Fabric | SNOVA
Host CPU PCIe POD (scale-up domain) Accelerator Accelerator Accelerator Accelerator Scale-up switch UALink fabric NIC border crossing: memory-semantic → network Leaf switch Ethernet / UEC
Figure 1. Pod-level scale-up domain over UALink with a border crossing to Ethernet/UEC for scale-out.
Executive Summary

AI training traffic splits cleanly into two classes: small-medium, latency-sensitive exchanges inside a pod (scale-up) and large, bandwidth-heavy exchanges across nodes (scale-out). UALink is built for the former; Ethernet/UEC for the latter. This division of labor simplifies microarchitecture, improves efficiency, and accelerates time-to-silicon.

The problem: one workload, two traffic classes

Modern AI training stacks generate both fine-grained, latency-critical messages (e.g., tensor-parallel collectives) within a pod and massive, throughput-oriented transfers (e.g., gradients) across nodes.

Trying to optimize a single interconnect for both leads to suboptimal designs and wasted power. Instead, we align each fabric with the traffic it serves best.

The two domains, precisely

Scale-up (intra-pod)

Connect accelerators and memory within a pod. Optimized for latency, ordering, and semantics.

Scale-out (inter-node)

Connect pods across a cluster. Optimized for bandwidth, reach, and congestion management.

Tensor-parallel exchanges
small-medium transfers, latency-sensitive
centimetres-metres
→
UALink
scale-up · memory-semantic · sub-μs latency
Gradient all-reduce / activations
large transfers, bandwidth-dominated
metres-kilometres
→
Ethernet / UEC
scale-out · packet transport · congestion-managed
Figure 2. Traffic classes mapped to the right fabric.
Characteristic UALink (scale-up) Ethernet UEC (scale-out)
Connectivity Intra-pod, accelerators & memory Inter-pod, rack/cluster
Semantics Memory-semantic, ordering, low jitter Packet transport, loss recoverable
Latency envelope Sub-μs μs-ms
Scale target Dozens of endpoints Thousands of endpoints
Physical layer Short reach, high efficiency Longer reach, higher aggregate BW
Failure model Fail-stop within pod Partial failures, congestion managed
Ecosystem state Emerging (UALink) Mature (Ethernet/UEC)
Figure 3. Key differences between UALink (scale-up) and Ethernet/UEC (scale-out).

What this means for silicon teams

1. Right tool, right job
Allocate SerDes lanes and PHY features to match each fabric's true requirements.
2. Clean boundaries reduce complexity
A clear border crossing (NIC/bridge) isolates semantics and simplifies verification.
3. Better power & performance
Short-reach, low-latency UALink inside the pod; bandwidth-optimized Ethernet outside.

Design & Verification Checklist

  • 1. Define scale-up/scale-out boundary in architecture spec — which traffic rides which fabric
  • 2. Budget SerDes lanes across UALink, Ethernet, and PCIe before RTL starts
  • 3. UALink: pin spec version, enumerate open choices, build protocol checkers and coverage
  • 4. Ethernet/UEC: treat FEC configuration and AN/LT as first-class verification targets
  • 5. Verify the seams: shared resets/clocks, error propagation, host-visible PCIe behavior
“” SNOVA perspective

Splitting the fabric is not about adding parts — it's about aligning each domain with the physics and semantics it serves best. We help teams define the boundary, allocate resources, and build verification plans that de-risk both fabrics.

— SNOVA Technologies

Conclusion

Scale-up and scale-out are different problems. Use UALink where latency and semantics matter most, and Ethernet/UEC where reach and bandwidth dominate. The result is a simpler architecture, lower power, and faster time-to-market.

Planning a scale-up architecture, or splitting SerDes between fabrics?
Your first conversation is with an engineer.
Discuss your scale-up architecture →