Skip to content
Fide Systems
Portfolio
Operating Founded 2022

Meridian Compute

Low-latency model serving across heterogeneous accelerators.

Sector
Inference infrastructure
Headquarters
Austin, TX
Ownership
Majority
Team
41

Overview

Meridian operates a serving fabric that places inference workloads across mixed fleets — current-generation GPUs, prior-generation cards still under lease, and dedicated accelerators — without the calling application knowing or caring which one answered.

Why this exists

Most teams buy inference capacity the way they bought servers in 2009: one vendor, one instance type, capacity sized to peak. The result is a fleet that is badly utilized at the median hour and short at the tail. Meanwhile the hardware under it turns over faster than any procurement cycle can absorb.

Meridian treats the fleet as one addressable pool. Requests are routed on live latency and cost per token rather than static assignment, and models are recompiled per target so that older silicon stays economically useful well past the point where a single-vendor deployment would have retired it.

Why we hold it

This is the layer that everything else in the group runs on, and it gets more valuable as accelerator supply gets more fragmented, not less. The switching cost for a customer who has moved production traffic onto the fabric is high and grows with every model they add.

Current focus

  • Deterministic tail latency under mixed-tenancy load
  • Speculative decoding across non-matching hardware pairs
  • Regional capacity commitments for customers with data-residency obligations