One cloud formodels, GPUs,and agents

Inference, compute, and agent runtimes behind a single API. Ship on day one.

TRUSTED BY

Halcyon
Tidepool
Kilobyte
Quorra
Openroute
Fishbowl
Hygo
Gizmo
Simular
Wizely
MODEL APIS
LLMImageAudioVideoVision
Model["BM-Sonnet-3"]
200+models
200mslatency
99.5%uptime
1Serverless Model APIs

Reach 200+ models through one endpoint. Nothing to provision.

Text, image, audio and video, all served on demand. Swap a model name and keep the same request shape. You are billed per token, never per idle hour.

Explore All Models
AAtlas 4 Pro$1.74/Mt Input ·
$3.48/Mt Output
1M Context
LLM
VVega M2$0.30/Mt Input ·
$1.20/Mt Output
200K Context
LLM
LLumen 5.1$1.40/Mt Input ·
$4.40/Mt Output
200K Context
LLM
CCorvid K2$0.95/Mt Input ·
$4.00/Mt Output
256K Context
LLM
EEmber 31B$0.14/Mt Input ·
$0.40/Mt Output
256K Context
LLM
SSable 397B$0.60/Mt Input ·
$3.60/Mt Output
256K Context
LLM
AAtlas 4 Pro$1.74/Mt Input ·
$3.48/Mt Output
1M Context
LLM
VVega M2$0.30/Mt Input ·
$1.20/Mt Output
200K Context
LLM
LLumen 5.1$1.40/Mt Input ·
$4.40/Mt Output
200K Context
LLM
CCorvid K2$0.95/Mt Input ·
$4.00/Mt Output
256K Context
LLM
EEmber 31B$0.14/Mt Input ·
$0.40/Mt Output
256K Context
LLM
SSable 397B$0.60/Mt Input ·
$3.60/Mt Output
256K Context
LLM
Base_url
["API.BLUEMETAL.DEV/YOUR-ENDPOINT"]
Operational
P99
53ms
Throughput
1,252 tok/s
Requests
19,421
Latency — last 40 requests
2Dedicated Endpoints

Private capacity. Predictable latency. No noisy neighbours.

Pin a model to compute that only you touch. Throughput stays flat as traffic climbs, so your p99 stops moving when it matters most.

Get Started
AGENT SANDBOX
Agent["CODING AGENTS"]
Coding agent · activeSandbox runtime
Run test suite · pytestQUEUED
Write fix · patch appliedRUNNING
Identify bug · null pointer line 84DONE
Read codebase · src/api/routes.pyDONE
startup
~200ms
isolation
full
billing
per second
status
running
1Agent sandbox

Isolated runtimes, built for agents that take real actions.

Not a notebook and not a container you have to babysit. A disposable environment where an agent can run commands, call tools and hit models — sealed off, every single time.

Get Started
GPU CLOUD
GPU[FLAGSHIP]
GPU model
NVIDIA H100Running
8× GPUs
GPU memory
80 GB HBM3×8
Instance specs
vCPU
96
Memory
960
Storage
2000
1GPU Instances

Full-control GPU machines, live in seconds.

Train, fine-tune or serve on hardware you own for the duration. Dedicated devices, root access, and performance that does not drift because of someone else's job.

Get Started
JobQueuedRunningComplete
Allocating GPU resourcesAllocating
12%
allocated
auto
duration
0.1s
cost
$0.0001
idle time
$0.00
2Serverless GPU

Hand us the job. We find the hardware.

Nothing to provision and nothing idling. Capacity scales up while the queue is deep and falls to zero the moment it drains — you pay for execution only.

Get Started
Cluster["CLUSTER-01"]
Cluster-01 · 6 nodesNVLink · GPUDirect RDMA · PCIe
Node-01
51%
Node-02
79%
Node-03
86%
Node-05
89%
Node-06
65%
Node-07
81%
GPU8× NVIDIA H200GPU memory141 GB HBM3e per GPU1.128 TB total
Nodes6 / 6InterconnectNVLink 4th Gen — 900 GB/sNetwork400 Gb/s RDMA
3Bare Metal

Peak throughput, with no hypervisor in the way.

Reserved physical clusters for large training runs and high-volume inference. NVLink inside the node, RDMA between them, and the whole machine answering to you.

Get Started
Why BlueMetal

Built for AI from day one. Designed for what you are actually shipping.

Get Started
Others
BlueMetal

Better price-performance

Up to 50% below the major clouds — the result of owning the infrastructure, not of cutting corners.

99.99%
Uptime SLA
<50ms
P50 latency
Global
PoPs
Operational

Built for production reliability

Low latency and high throughput that hold their shape under sustained, real-world load.

Agent runtimeSANDBOX
Model APIs200+ MODELS
GPU infrastructureBARE METAL

One platform for the full AI stack

Model APIs, GPU capacity and agent runtimes sit behind a single account and a single bill.

Serverless API
Dedicated endpoints
GPU control

Scale with your workload

Begin on the shared API and graduate to reserved clusters without rewriting a line.

Seeing latency spikes in prod — can you help debug?2 MIN AGOOn it — pulling your logs now. Config issue, fix incoming.JUST NOW

Dedicated support when it matters

Direct access to engineers who run this infrastructure themselves, not a ticket queue.

Testimonials

Don’t take our word for it.

Halcyon

New checkpoints show up on BlueMetal within a day of release, which means our users get them before we have finished reading the paper. That cadence is hard to find anywhere else.

MRM. RiveraCo-Founder & CTO
Gizmo

Integration took an afternoon. The API behaves the same whether we send ten requests or ten thousand, so scaling our study tools has been a non-event.

PCP. CastellanosCo-Founder and CEO
Fishbowl

Our speech models sit on BlueMetal GPUs and simply stay up. We spend our time on the model instead of on capacity planning and driver versions.

SLS. LindqvistCo-Founder & Chief Scientist
Kilobyte

Agentic coding punishes slow inference. BlueMetal gives us consistent throughput across several models at once, and the team tunes for the workloads we actually run.

AMA. MoreauHead of Partnerships

Everything you need to run AI in production.

Hundreds of models, on-demand GPUs and sealed agent runtimes, unified behind one API. Free to start, priced to keep going.

Get Started