Run vLLM at scale without operating inference infrastructure.
Set SLO and cost targets. Rivvr runs and continuously optimizes your infrastructure to meet them,
inside your AWS account.
Up to 70% cost savings
SLOs met automatically
Bring your own models
Built by engineers from:




With SageMaker or K8s
Configure infrastrucure, routing, scaling, select GPUs
Tune infrastrucure and load test vLLM after each model, LoRA or harness update
Overprovision to protect SLOs during traffic spikes
Investigate failures and restore service
With Rivvr
Define your model, SLOs, and cost targets
Rivvr automatically load tests, tunes, and deploys your model
Rivvr automatically meets SLOs during 10× traffic spikes while minimizing costs
Rivvr detects failures and restores serving capacity automatically
How is this different from SageMaker?
SageMaker gives you infrastructure and autoscaling controls. Rivvr removes the need to operate them: set latency and cost targets, and Rivvr continuously runs the infrastructure for you.
How is this different from AIBrix or NVIDIA Dynamo?
AIBrix and Dynamo are infrastructure software your team deploys, configures and maintains. Rivvr is a managed platform inside your AWS account. Instead of configuring profiling, autoscaling and infrastructure parameters, you specify the outcome: SLO + cost target.
Isn't spot capacity risky for something with an SLO?
Spot capacity is interruptible. Rivvr spreads workloads across compatible GPU capacity pools and uses on-demand as fallback, allowing capacity to change without sacrificing the SLO.
What data leaves my AWS account?
Inside your AWS account, behind your VPC. Inference traffic, data, and model weights never leave your environment.
How is Rivvr priced?
You pay AWS directly for infrastructure. Rivvr charges a separate management fee based on GPUs under orchestration.
Do I need to change my application?
Usually not. Rivvr exposes an OpenAI-compatible API, so most teams only change the endpoint.
Does Rivvr modify my model?
No. Model weights remain unchanged. Rivvr optimizes orchestration, placement and infrastructure.
What models are supported?
Any LLM supported by vLLM, currently up to 400B parameters.
What GPUs are supported?
NVIDIA CUDA compute capability 7.0+, including L4, L40S, A10G, T4, A100, H100, H200, B200, B300 and V100.
