AI training traffic splits cleanly into two classes: small-medium, latency-sensitive exchanges inside a pod (scale-up) and large, bandwidth-heavy exchanges across nodes (scale-out). UALink is built for the former; Ethernet/UEC for the latter. This division of labor simplifies microarchitecture, improves efficiency, and accelerates time-to-silicon.
The problem: one workload, two traffic classes
Modern AI training stacks generate both fine-grained, latency-critical messages (e.g., tensor-parallel collectives) within a pod and massive, throughput-oriented transfers (e.g., gradients) across nodes.
Trying to optimize a single interconnect for both leads to suboptimal designs and wasted power. Instead, we align each fabric with the traffic it serves best.
The two domains, precisely
Scale-up (intra-pod)
Connect accelerators and memory within a pod. Optimized for latency, ordering, and semantics.
Scale-out (inter-node)
Connect pods across a cluster. Optimized for bandwidth, reach, and congestion management.
| Characteristic | UALink (scale-up) | Ethernet UEC (scale-out) |
|---|---|---|
| Connectivity | Intra-pod, accelerators & memory | Inter-pod, rack/cluster |
| Semantics | Memory-semantic, ordering, low jitter | Packet transport, loss recoverable |
| Latency envelope | Sub-μs | μs-ms |
| Scale target | Dozens of endpoints | Thousands of endpoints |
| Physical layer | Short reach, high efficiency | Longer reach, higher aggregate BW |
| Failure model | Fail-stop within pod | Partial failures, congestion managed |
| Ecosystem state | Emerging (UALink) | Mature (Ethernet/UEC) |
What this means for silicon teams
Design & Verification Checklist
-
1. Define scale-up/scale-out boundary in architecture spec — which traffic rides which fabric
-
2. Budget SerDes lanes across UALink, Ethernet, and PCIe before RTL starts
-
3. UALink: pin spec version, enumerate open choices, build protocol checkers and coverage
-
4. Ethernet/UEC: treat FEC configuration and AN/LT as first-class verification targets
-
5. Verify the seams: shared resets/clocks, error propagation, host-visible PCIe behavior
Splitting the fabric is not about adding parts — it's about aligning each domain with the physics and semantics it serves best. We help teams define the boundary, allocate resources, and build verification plans that de-risk both fabrics.
Conclusion
Scale-up and scale-out are different problems. Use UALink where latency and semantics matter most, and Ethernet/UEC where reach and bandwidth dominate. The result is a simpler architecture, lower power, and faster time-to-market.