Zero-Ops LLM Inference
Inside Your AWS

Zero-Ops LLM Inference
Inside Your AWS

Run vLLM at scale without operating inference infrastructure.
Set SLO and cost targets. Rivvr runs and continuously optimizes your infrastructure to meet them,
inside your AWS account.

Up to 70% cost savings

SLOs met automatically

Bring your own models

VPC
Model + LoRAs
SLO + cost targets
Rivvr Orchestrator
SLO controller
Cost optimizer
Load tester
Performance monitor
Mixed GPU Fleet
H100
Spot
vLLM · GPT-OSS 20B
L40S
Spot
vLLM · GPT-OSS 20B
A100
On-demand
vLLM · GPT-OSS 20B
L4
Spot
vLLM · GPT-OSS 20B
Live metrics

Built by engineers from:

University of Toronto logo
SickKids Research Institute logo
Xiaomi logo
OneShield logo

Stop managing inference infrastructure

Stop managing inference infrastructure

With SageMaker or K8s

Configure infrastrucure, routing, scaling, select GPUs

Tune infrastrucure and load test vLLM after each model, LoRA or harness update

Overprovision to protect SLOs during traffic spikes

Investigate failures and restore service

With Rivvr

Define your model, SLOs, and cost targets

Rivvr automatically load tests, tunes, and deploys your model

Rivvr automatically meets SLOs during 10× traffic spikes while minimizing costs

Rivvr detects failures and restores serving capacity automatically

You set the targets.
Rivvr runs the inference.

You set the targets.
Rivvr runs the inference.

Define requirements, not your infrastructure

Bring your model weights from Hugging Face, S3, or your local store. Set your latency, throughput, and cost targets. Use the defaults, or customize context length and quantization.

Define requirements, not your infrastructure

Bring your model weights from Hugging Face, S3, or your local store. Set your latency, throughput, and cost targets. Use the defaults, or customize context length and quantization.

Deployment Policy
llama-3-70b
Quantization
LoRa
TTFT p50
Output Token/s
TTFT p90
Cost guardrail
Speed mode
Loading...

Rivvr automatically meets your SLOs

Rivvr continuously optimizes your serving stack, including GPU kernels, KV cache, routing, and orchestration, to meet SLOs during 10× traffic spikes while minimizing costs. Orchestrator monitors production, isolates failures with circuit breakers, and automatically restores serving capacity even under load.

Rivvr automatically meets your SLOs

Rivvr continuously optimizes your serving stack, including GPU kernels, KV cache, routing, and orchestration, to meet SLOs during 10× traffic spikes while minimizing costs. Orchestrator monitors production, isolates failures with circuit breakers, and automatically restores serving capacity even under load.

req / s
llama-3-70b
req / s
mistral-7b
GPU Cluster
A100
80GB
on-demandspot
llama-3-70b
65%
H100
80GB
spot
llama-3-70b
70%
L40S
48GB
spoton-demand
llama-3-70b
58%
L4
24GB
spot
mistral-7bllama-3-70b
62%
A100
40GB
on-demandspot
mistral-7bllama-3-70b
55%
L4
24GB
spoton-demand
mistral-7b
60%

Observe performance. Change your targets.

Track live latency, throughput, and costs across your models. Tighten a latency target or change a cost limit. Rivvr adjusts the serving stack without manual redeployment or downtime.

Observe performance. Change your targets.

Track live latency, throughput, and costs across your models. Tighten a latency target or change a cost limit. Rivvr adjusts the serving stack without manual redeployment or downtime.

Live Monitoring
llama-3-70b
TTFT
p50 0msp90 0ms
Tokens / s
p50 0p90 0
Prefill / M tokens
$0.00
Decode / M tokens
$0.00
SLO Compliance0.0%

Built for production workloads

Built for production workloads

Voice AI

Support chat agents

Coding assistants & agents

OCR & document intake

THROUGHPUT-BOUND

The Challenge

Document batches create uneven GPU demand, with spikes followed by idle periods.

Shared infrastructure bills hide which models drive inference spend.

With Rivvr

Automatically adjust GPU capacity as document batches arrive and complete.

Track inference spend by model and cost per token.

Voice AI

Support chat agents

Coding assistants & agents

OCR & document intake

THROUGHPUT-BOUND

The Challenge

Document batches create uneven GPU demand, with spikes followed by idle periods.

Shared infrastructure bills hide which models drive inference spend.

With Rivvr

Automatically adjust GPU capacity as document batches arrive and complete.

Track inference spend by model and cost per token.

Voice AI

Support chat agents

Coding assistants & agents

OCR & document intake

LATENCY-CRITICAL

The Challenge

Unpredictable LLM response times create awkward pauses in conversations.

Keeping GPUs ready for peak call volume means paying for idle capacity.

With Rivvr

Maintain first-token latency targets as concurrent calls increase.

Automatically adjust GPU capacity to call volume while reducing idle spend.

Runs in your cloud.
Stays in your cloud.

Runs in your cloud.
Stays in your cloud.

VPC deployment

Run inference inside your AWS account and VPC, with network controls you manage.

VPC deployment

Run inference inside your AWS account and VPC, with network controls you manage.

Data privacy

Your prompts, responses, logs, and model weights stay inside your AWS account.

Data privacy

Your prompts, responses, logs, and model weights stay inside your AWS account.

RBAC and access control

Control who can deploy models, change policies, and view cost and performance data.

RBAC and access control

Control who can deploy models, change policies, and view cost and performance data.

Flexible tenancy isolation

Choose dedicated infrastructure for strict isolation or shared GPU capacity for higher utilization.

Flexible tenancy isolation

Choose dedicated infrastructure for strict isolation or shared GPU capacity for higher utilization.

Everything to run inference in production

Everything to run inference in production

Bring your own weights

Deploy vLLM-supported models from Hugging Face, S3, or Rivvr model storage. Model weights stay inside your AWS account.

Managed endpoints

Get a managed, OpenAI-compatible endpoint for real-time and batch inference.

Multi-model and multi-LoRA serving

Share GPU capacity across multiple models and LoRA adapters, with allocation that adapts to demand.

Automatic load testing

Find efficient vLLM configurations through load tests across GPU types. Retest automatically as workloads and SLO targets change.

Bring your own weights

Deploy vLLM-supported models from Hugging Face, S3, or Rivvr model storage. Model weights stay inside your AWS account.

Managed endpoints

Get a managed, OpenAI-compatible endpoint for real-time and batch inference.

Multi-model and multi-LoRA

Share GPU capacity across multiple models and LoRA adapters, with allocation that adapts to demand.

Automatic load testing

Find efficient vLLM configurations through load tests across GPU types. Retest automatically as workloads and SLO targets change.

Frequently Asked Questions

Frequently Asked Questions

How is this different from SageMaker?

SageMaker gives you infrastructure and autoscaling controls. Rivvr removes the need to operate them: set latency and cost targets, and Rivvr continuously runs the infrastructure for you.

How is this different from AIBrix or NVIDIA Dynamo?

AIBrix and Dynamo are infrastructure software your team deploys, configures and maintains. Rivvr is a managed platform inside your AWS account. Instead of configuring profiling, autoscaling and infrastructure parameters, you specify the outcome: SLO + cost target.

Isn't spot capacity risky for something with an SLO?

Spot capacity is interruptible. Rivvr spreads workloads across compatible GPU capacity pools and uses on-demand as fallback, allowing capacity to change without sacrificing the SLO.

What data leaves my AWS account?

Inside your AWS account, behind your VPC. Inference traffic, data, and model weights never leave your environment.

How is Rivvr priced?

You pay AWS directly for infrastructure. Rivvr charges a separate management fee based on GPUs under orchestration.

Do I need to change my application?

Usually not. Rivvr exposes an OpenAI-compatible API, so most teams only change the endpoint.

Does Rivvr modify my model?

No. Model weights remain unchanged. Rivvr optimizes orchestration, placement and infrastructure.

What models are supported?

Any LLM supported by vLLM, currently up to 400B parameters.

What GPUs are supported?

NVIDIA CUDA compute capability 7.0+, including L4, L40S, A10G, T4, A100, H100, H200, B200, B300 and V100.

See it running on your setup

See it running on your setup