Zero inference ops
Meet latency SLOs automatically
Run any vLLM model or LoRA
With SageMaker or DIY
Configure the infrastructure
Choose GPUs, capacity, and scaling
Retune as traffic changes
Monitor and optimize, manually
With Rivvr
Set your latency and cost targets
Rivvr chooses the infrastructure
Adapts automatically as traffic changes
Continuously optimized, in the background
How is this different from SageMaker?
SageMaker gives you the knobs for instance sizing and autoscaling. Someone still has to turn and maintain them. Rivvr means you don't need that person: set your latency and cost target, and the infrastructure runs itself.
How is this different from AIBrix or NVIDIA Dynamo?
Two things. First: you don't operate anything — Rivvr runs as a managed platform inside your AWS account, unlike the software your team deploys and maintains. Second: a different operational model. AIBrix and Dynamo ask you to configure autoscaling, profiling, and dozens of other parameters. Rivvr asks for two — your SLO and your cost guardrail. Set those, and you're running.
Isn't spot capacity risky for something with an SLO?
Alone, yes — spot can be reclaimed with little warning. Rivvr's cluster spans multiple GPU types, so there's rarely just one spot pool to lose — if one type is reclaimed, another picks up the load, with on-demand as the fallback. More pools, more savings, no interruption risk.
Where does Rivvr run?
Inside your AWS account, behind your VPC. Inference traffic, data, and model weights never leave your environment.
How is Rivvr priced?
Rivvr is priced as a management layer based on GPUs under orchestration. You pay AWS directly for infrastructure; Rivvr charges a separate management fee.
Do I need to change my integration code?
No. Rivvr uses an OpenAI-compatible API. Most teams point their existing client at a new endpoint, and nothing else changes.
Does Rivvr modify or compress my models?
No. Rivvr doesn't touch model weights. Optimization happens at the orchestration, placement, and infrastructure layers.
What models does Rivvr support?
Any Large Language Model served by vLLM, up to 400B parameters. If you're running something larger or unusual, let's talk.
What inference engines are supported?
Rivvr uses a vLLM-compatible runtime.
What GPUs are supported?
NVIDIA GPUs with CUDA compute capability 7.0 or higher — including L4, L40S, A10G, T4, A100, H100, H200, B200, B300, V100.
