INFERENCE API

GLM-5.3 is live.

33 models on one OpenAI-compatible endpoint, one key, up to 1M context.

Everything you need to
build and scale.

GPU Pods

GPU Pods

H100, H200 and B200 accelerators on demand. Reserve a single GPU or a multi-node cluster, provisioned in under 90 seconds and billed by the second, with persistent volumes that follow your workload between sessions.

H100H200B200B300
Explore GPU Pods

Need bigger?
Reserve a cluster.

Multi-node H200, B200 and B300 clusters with NVIDIA NVLink 5.0 fabric, dedicated capacity and committed pricing. From a single 8-GPU node to thousand-GPU training runs; we handle the rest.

Ready to build?
Talk to our infrastructure team and get a custom quote.
Multi-node NVLink fabric
8× B300 SXM per node · 1.8 TB/s GPU-to-GPU
Reserved pricing
Up to 60% off on-demand · 1-mo to 3-yr terms
Dedicated support
Priority access, 24/7 coverage, SLA-backed
Rapid provisioning
Clusters ready in hours, not days
Custom networking
Tailored VPC, routing and isolation
99.99% uptime SLA
Enterprise-grade reliability you can build on
Cluster configurationReview & request
NVIDIA B300 SXM · 8-node cluster
NVLink fabric · Redundant power · PCIe 5.0
GPUs
64× NVIDIA B300
GPU memory
18.4 TB (288 GB / GPU)
Interconnect
NVLink Switch System 5.0
vCPUs / node
96 vCPUs
Networking
800 Gbps · RDMA
Term
12-month reserved
Request this configuration

Every provider behind
a single integration.

33 models, 8 providers, one OpenAI- and Anthropic-compatible endpoint. Switch providers with a string change.

33
Models
8
Providers
1M
Max context
1
API key
StreamingTool callingJSON modeVisionLong contextBYOKOff-peak pricing
Browse the full catalog

claude-opus-5

Anthropic

Frontier · Tools · 1M context

gpt-6-astra

OpenAI

Frontier · Tools · 1M context

grok-4.6

xAI

Reasoning · Tools · 500k

kimi-k3

Moonshot

Open weights · Tools · 1M

glm-5.3

Zhipu

Open weights · Tools · 1M

deepseek-v4-pro

DeepSeek

Reasoning · Tools · 1M

claude-opus-5

Anthropic

Frontier · Tools · 1M context

gpt-6-astra

OpenAI

Frontier · Tools · 1M context

grok-4.6

xAI

Reasoning · Tools · 500k

kimi-k3

Moonshot

Open weights · Tools · 1M

glm-5.3

Zhipu

Open weights · Tools · 1M

deepseek-v4-pro

DeepSeek

Reasoning · Tools · 1M

claude-sonnet-5

Anthropic

Balanced · Tools · 1M context

gpt-5.6-terra

OpenAI

Balanced · Tools · 1M context

grok-4.5

xAI

Reasoning · Tools · 500k

kimi-k2.7-code

Moonshot

Code · Open weights · 256k

minimax-m3

MiniMax

Open weights · Tools · 1M

glm-5.3-flash-derisked

Zhipu

Derisked · Hosted · 1M

claude-sonnet-5

Anthropic

Balanced · Tools · 1M context

gpt-5.6-terra

OpenAI

Balanced · Tools · 1M context

grok-4.5

xAI

Reasoning · Tools · 500k

kimi-k2.7-code

Moonshot

Code · Open weights · 256k

minimax-m3

MiniMax

Open weights · Tools · 1M

glm-5.3-flash-derisked

Zhipu

Derisked · Hosted · 1M

claude-haiku-4.5

Anthropic

Fast · Tools · 200k

gpt-5.4-mini

OpenAI

Fast · Tools · 400k

gpt-5.3-codex

OpenAI

Code · Tools · 400k

deepseek-v4-flash

DeepSeek

Fast · Open weights · 1M

doubao-seed-2.1-turbo

ByteDance

Fast · Tools · 256k

claude-fable-5.1

Anthropic

Frontier · Tools · 1M context

claude-haiku-4.5

Anthropic

Fast · Tools · 200k

gpt-5.4-mini

OpenAI

Fast · Tools · 400k

gpt-5.3-codex

OpenAI

Code · Tools · 400k

deepseek-v4-flash

DeepSeek

Fast · Open weights · 1M

doubao-seed-2.1-turbo

ByteDance

Fast · Tools · 256k

claude-fable-5.1

Anthropic

Frontier · Tools · 1M context

Compute

Three tiers on one control plane. Move between them without rebuilding your stack: same API, same billing, same audit trail, whether you are running a staging box or a regulated workload on its own hardware.

15 regions · hourly + monthly · API-first
01

Shared

Virtual machines on shared vCPUs. Dev, staging and small production.

From
$5
/ month · billed hourly
  • 1 to 32 vCPU, 1 GB to 192 GB RAM
  • 25 GB to 3.8 TB NVMe storage
  • Full root access, SSH keys
  • Snapshots and backups
  • Billed by the hour
Most popular
02

VDS

Virtual dedicated servers on pinned physical cores. Production APIs and databases.

From
$24
/ month · billed hourly
  • Dedicated cores, no noisy neighbours
  • 2 to 256 vCPU, 4 GB to 512 GB RAM
  • 41 GB to 5.1 TB NVMe storage
  • Guaranteed baseline performance
  • Full root access, SSH keys
  • Same API and billing as Shared
03

Bare metal servers

A whole physical server, no hypervisor. Your kernel, your hardware.

From
$59
/ month · 16 configurations
  • 6 to 192 cores, 64 GB to 1.5 TB RAM
  • AMD EPYC and Ryzen, Intel Xeon
  • IPMI out-of-band access
  • Custom OS, hardware RAID
  • Single tenant, monthly term

Global infrastructure

15 regions·13 countries·In-region residency
world map
MumbaiSingaporeDubaiTokyoSydneyLondonFrankfurtAmsterdamParisMadridStockholmNew YorkSan FranciscoLos AngelesSão Paulo
Ready to build?

Infrastructure that moves
at the speed of your ideas.