Managed LLM Inference
Inside Your AWS

Managed LLM Inference Inside Your AWS

Run vLLM at scale without operating inference infrastructure.
Set SLO and cost targets. Rivvr handles infrastructure to meet them,
inside your AWS account.

Run vLLM at scale without operating inference infrastructure.
Set SLO and cost targets.
Rivvr handles infrastructure to meet them, inside your AWS account.

Up to 70% cost savings

Zero inference ops

Any vLLM model or LoRA

VPC
Model + LoRAs
SLO + cost targets
Rivvr Orchestrator
SLO control
Cost optimization
KV management
Observability
Mixed GPU Fleet
H100
Spot
vLLM · GPT-OSS 20B
L40S
Spot
vLLM · GPT-OSS 20B
A100
On-demand
vLLM · GPT-OSS 20B
L4
Spot
vLLM · GPT-OSS 20B
Live metrics

Built by engineers from:

University of Toronto logo
SickKids Research Institute logo
Xiaomi logo
OneShield logo

Stop configuring inference infrastructure

Stop configuring inference infrastructure

With SageMaker or DIY

Configure serving infrastructure

Load test and tune each model

Monitor production and scaling

Retune as models and traffic change

Recover failures and maintain the stack

With Rivvr

Define model + performance requirements

Rivvr profiles and deploys it

Rivvr monitors production continuously

Rivvr adapts as models and traffic change

Rivvr recovers and keeps it running

With Rivvr

Define model + performance requirements

Rivvr profiles and deploys it

Rivvr monitors production continuously

Continuously optimized, in the background

Rivvr recovers and keeps it running

You set the targets.
Rivvr runs the inference.

You set the targets.
Rivvr runs the inference.

Define your SLOs, not your infrastructure

Bring weights from Hugging Face, S3, or your local store. Define desired context length, model and KV cache quantization, or leave defaults. Set application requirements: TTFT, TPS, p50/p90, cost guardrails and SLO complience level.

Define your SLOs, not your infrastructure

Bring your own weights, or import from Hugging Face. Define model configs — data type, quantization, KV cache quantization, context length — or leave defaults. Set TTFT, TPS, p50/p90, cost guardrails, and speed mode.

Deployment Policy
llama-3-70b
Quantization
LoRa
TTFT p50
Output Token/s
TTFT p90
Cost guardrail
Speed mode
Loading...

Rivvr continuously operates your inference

Rivvr continusly re-tests and tunes vLLM as models or traffic change, monitors production performance, re-optimizes the mixed-GPU fleet, and recovers from infrastructure failures. Zero operations, up to 70% cost savings.

Rivvr continuously operates your inference

Rivvr continusly re-tests and tunes vLLM as models or traffic change, monitors production performance, re-optimizes the mixed-GPU fleet, and recovers from infrastructure failures. Zero operations, up to 70% cost savings.

req / s
llama-3-70b
req / s
mistral-7b
GPU Cluster
A100
80GB
on-demandspot
llama-3-70b
65%
H100
80GB
spot
llama-3-70b
70%
L40S
48GB
spoton-demand
llama-3-70b
58%
L4
24GB
spot
mistral-7bllama-3-70b
62%
A100
40GB
on-demandspot
mistral-7bllama-3-70b
55%
L4
24GB
spoton-demand
mistral-7b
60%

Observe performance. Change requirements.

Track live performance and costs of your model fleet. Tighten a latency target, change a cost guardrail. Rivvr automatically re-adjusts the serving stack without redeploy or downtime.

Observe performance. Change requirements.

Track live TTFT, TPS, SLO compliance, cost per model, cost per token. Change your targets anytime — no redeploy, no downtime. Tighten an SLO, raise a cost guardrail, see the cluster adapt.

Live Monitoring
llama-3-70b
TTFT
p50 0msp90 0ms
Tokens / s
p50 0p90 0
Prefill / M tokens
$0.00
Decode / M tokens
$0.00
SLO Compliance0.0%

Built for production
AI agents

Built for production
AI agents

Voice AI

Support chat agents

Coding assistants & agents

OCR & document intake

LATENCY-CRITICAL

The Challenge

Any TTFT variance is an audible pause; there’s no graceful degradation

Static warm pools protect latency, but bill you for capacity you don’t use

With Rivvr

Your pipeline stays under 400ms without a dedicated warm pool

GPU spend tracks actual call volume, not your peak estimate

Voice AI

Support chat agents

Coding assistants & agents

OCR & document intake

LATENCY-CRITICAL

The Challenge

Any TTFT variance is an audible pause; there’s no graceful degradation

Static warm pools protect latency, but bill you for capacity you don’t use

With Rivvr

Your pipeline stays under 400ms without a dedicated warm pool

GPU spend tracks actual call volume, not your peak estimate

Runs in your cloud.
Stays in your cloud.

Runs in your cloud.
Stays in your cloud.

VPC deployment

Runs inside your AWS account, in a network-isolated environment you control.

VPC deployment

Runs inside your AWS account, in a network-isolated environment you control.

Data privacy

Prompts, responses, logs, and model weights stay in your cloud.

Data privacy

Prompts, responses, logs, and model weights stay in your cloud.

RBAC and access control

RBAC controls who can deploy models, change policies, and view cost or performance data.

RBAC and access control

RBAC controls who can deploy models, change policies, and view cost or performance data.

Flexible tenancy isolation

Use dedicated infrastructure where strict isolation is required or shared capacity where efficiency matters.

Flexible tenancy isolation

Use dedicated infrastructure where strict isolation is required or shared capacity where efficiency matters.

Everything to run inference in production

Everything to run inference in production

Spot + on-demand

Rivvr automatically mixes lower-cost spot pool with reliable on-demand, keeping deployment SLO-complient at loweest cost possible.

Heterogenious orchestration

Heterogenious-first planner that automatically chooses the best GPU type for each model and workload to meet SLO and optimize costs.

Fast cold starts

Bring new model capacity online within 10 seconds so the fleet can react quickly to bursts and recover from failure.

KV Cache Optimization

Full agentic optimizations stack: intelligent routing, SLO-aware management, distributed KV cache across RAM, NVME, Redis and persistent store.

Bring your own weights

Deploy any vLLM-supported model from Hugging Face, S3, or Rivvr model storage. Keep model weights inside your AWS.

Managed endpoints

Every deployment gets a production-ready, OpenAI-compatible endpoint for real-time and batch processing.

Multi-model & Multi-LoRA

Run multiple models and LoRA adapters dynamically sharing GPU compute.

Automatic load testing

Profile models on heterogenious GPU fleet. Find most efficient vLLM configs for SLO targets. Re-test after traffic, model or SLO changes.

Frequently Asked Questions

Frequently
Asked Questions

How is this different from SageMaker?

SageMaker gives you infrastructure and autoscaling controls. Rivvr removes the need to operate them: set latency and cost targets, and Rivvr continuously runs the infrastructure for you.

How is this different from AIBrix or NVIDIA Dynamo?

AIBrix and Dynamo are infrastructure software your team deploys, configures and maintains. Rivvr is a managed platform inside your AWS account. Instead of configuring profiling, autoscaling and infrastructure parameters, you specify the outcome: SLO + cost target.

Isn't spot capacity risky for something with an SLO?

Spot capacity is interruptible. Rivvr spreads workloads across compatible GPU capacity pools and uses on-demand as fallback, allowing capacity to change without sacrificing the SLO.

Where does Rivvr run?

Inside your AWS account, behind your VPC. Inference traffic, data, and model weights never leave your environment.

How is Rivvr priced?

You pay AWS directly for infrastructure. Rivvr charges a separate management fee based on GPUs under orchestration.

Do I need to change my application?

Usually not. Rivvr exposes an OpenAI-compatible API, so most teams only change the endpoint.

Does Rivvr modify my model?

No. Model weights remain unchanged. Rivvr optimizes orchestration, placement and infrastructure.

What models are supported?

Any LLM supported by vLLM, currently up to 400B parameters.

What runtime does Rivvr use?

A vLLM-compatible runtime.

What GPUs are supported?

NVIDIA CUDA compute capability 7.0+, including L4, L40S, A10G, T4, A100, H100, H200, B200, B300 and V100.

See it running on your setup

See it running on your setup