Home Operated AI Inference Token API Platform GPU Rental Services About Us Contact Us Get a Quote
GPU rental & operated inference live today

Rent GPU Servers or
Skip the Ops Entirely

LEOMO Company operates NVIDIA, AMD and Intel bare-metal infrastructure. Rent the hardware by the hour, or call our OpenAI-compatible Token API and pay only for what you generate.

GPU per Node
2.3 TB
Max HBM per Node
400G
InfiniBand Fabric
< 200ms
API First Token

GPU Server Rental

Bare-metal NVIDIA, AMD and Intel servers by the hour or reserved monthly. Full root access, no hypervisor.

See rental plans

Operated AI Inference

Choose your model. We run it on dedicated GPUs with autoscaling, observability and data residency controls built in.

Explore operated inference

Token API Platform

OpenAI-compatible endpoint, billed per token. No GPUs to manage, no idle capacity to pay for.

View token pricing
NVIDIA BlackwellNVIDIA HopperAMD Instinct Intel GaudiOpenAI-Compatible API400G InfiniBand
GPU Rental Fleet

Available Accelerators & Rental Rates

Filter by manufacturer to compare specifications, on-demand hourly rates and reserved monthly rates. Pricing reflects current US and Japan cloud market benchmarks.

Showing all 8 GPU accelerators
NVIDIA In Stock
NVIDIABlackwell Ultra

B300

Frontier training and large-scale inference with 288GB HBM3e and 2× Hopper FP4 throughput.

  • Memory288GB HBM3e
  • Bandwidth8.2 TB/s
  • FP42× Hopper
  • ConfigHGX / SXM
On-Demand $13.20/GPU/hr
Reserved $9.50/GPU/hr
NVIDIA In Stock
NVIDIAHopper

H200

141GB HBM3e delivers up to 1.9× faster LLM inference than H100 at competitive cost.

  • Memory141GB HBM3e
  • Bandwidth4.8 TB/s
  • FP84 PFLOPS
  • ConfigSXM / NVL
On-Demand $6.31/GPU/hr
Reserved $5.40/GPU/hr
NVIDIA In Stock
NVIDIAHopper

H100 SXM

The industry-standard AI training accelerator, supported by every major framework and cloud toolchain.

  • Memory80GB HBM3
  • Bandwidth3.35 TB/s
  • FP84 PFLOPS
  • NVLink900 GB/s
On-Demand $6.16/GPU/hr
Reserved $4.50/GPU/hr
NVIDIA In Stock
NVIDIAAda Lovelace

L40S

48GB GDDR6 with strong FP16/INT8 throughput — the best price-performance entry point for AI inference.

  • Memory48GB GDDR6
  • Bandwidth864 GB/s
  • Form FactorPCIe Gen4
  • Best ForInference / Rendering
On-Demand $2.56/GPU/hr
Reserved $2.25/GPU/hr
NVIDIA In Stock
NVIDIAAmpere

A100 SXM

80GB HBM2e with proven reliability — still the workhorse for fine-tuning and mid-scale training.

  • Memory80GB HBM2e
  • Bandwidth2.0 TB/s
  • FP16312 TFLOPS
  • NVLink600 GB/s
On-Demand $3.67/GPU/hr
Reserved $2.70/GPU/hr
NVIDIA In Stock
NVIDIABlackwell

RTX PRO 6000

96GB GDDR7 workstation GPU — the most cost-effective entry point for AI, rendering and visualisation.

  • Memory96GB GDDR7
  • ArchitectureBlackwell
  • Form FactorPCIe Gen5
  • Best ForAI / Rendering
On-Demand $2.50/GPU/hr
Reserved $2.19/GPU/hr
Rental Plans

Rent by the Hour or Reserve for Less

No capital expenditure, no hardware to maintain. Choose on-demand flexibility or lock in reserved capacity at up to 40% below list rate.

On-Demand Rental

Hourly billing, deploy in minutes, terminate anytime. Ideal for experiments and burst workloads.

$0.99/GPU/hr
Starting from · L40S
  • No commitment, cancel anytime
  • Full root access to bare metal
  • 400G InfiniBand available
  • 24/7 technical support
Start Renting

Enterprise Cluster

Multi-node deployments with NVLink, Infinity Fabric and 400G InfiniBand, validated for distributed training.

Custom
Quoted per deployment
  • 8× GPU nodes and above
  • Fabric cabling & cluster validation
  • NCCL / RCCL tuning included
  • Onsite or remote handover
Design a Cluster
Business 02 · Operated AI Inference

Choose the Model. We Run It.

Deploy any open, commercial or custom model on managed inference infrastructure. Autoscaling, observability and data residency controls are built in — you never touch a GPU.

Production inference without the platform team

Run an open-source model, a commercial model, or your own fine-tuned weights through a managed endpoint. LEOMO operates the serving environment, handles autoscaling, and keeps your application interface stable while models and infrastructure evolve underneath.

  • Bring any model. Open-source catalog, commercial model endpoints, or your own private checkpoint. Define it with a config file or a custom container — we deploy it on dedicated GPUs.
  • Autoscaling with scale-to-zero. Set replica thresholds to match your traffic. Scale down to zero when idle to eliminate compute costs, scale up automatically when requests arrive.
  • Dedicated, predictable performance. Your model runs on dedicated GPUs with no rate limits and no noisy-neighbour latency. Throughput is limited only by your deployment’s capacity.
  • Data residency & retention controls. Choose where your inference runs. Zero data retention is available for eligible deployments; prompts and completions are never used for training.
  • Deep observability. Per-endpoint latency, throughput, error rates and GPU utilisation — with log export to your own monitoring stack.
  • 99.9% availability SLA. Same SLA as our dedicated rental plans, with 4-hour hardware replacement and 24/7 engineer access.
Operated Inference · Deployment Options
Model source Open · Commercial · Custom
Deployment method Config file or custom container
GPU allocation Dedicated · single-tenant
Autoscaling Replica-based · scale-to-zero
Endpoint OpenAI-compatible REST
Placement Hosted or on-premises
Data retention Zero-retention option
Availability SLA 99.9%
Billing Reserved GPU capacity
Business 03 · Token API Platform

One API Key. Pay Only for Tokens.

Call 100+ models through a shared per-token API. No provisioning, no replicas to size, no minimum cost — and no idle GPUs to pay for.

Serverless inference, OpenAI-compatible

Change one base URL and your existing OpenAI SDK code keeps working. The Token API is the fastest way to run inference on LEOMO hardware — you are billed on input and output tokens, with no capacity to reserve and no commitment.

  • 100+ models, one endpoint. Chat, reasoning, embedding and vision models from every major open-source family — all behind https://api.leomo.com/v1.
  • Input / output / cached token billing. You pay per million tokens processed. Cached input tokens are billed at a steep discount when the same prompt prefix repeats.
  • Streaming, tools & structured outputs. SSE streaming, function calling, JSON mode and vision inputs are supported across the catalog.
  • Batch inference at 50% off. Submit asynchronous batch jobs for non-latency-sensitive workloads and pay half the serverless rate on both input and output tokens.
  • Scoped API keys. Per-key rate limits, spend caps and usage attribution — no shared-key surprises, no runaway bills.
  • No lock-in. Export logs, usage records and evaluation results at any time. Swap models with a single parameter change.
api.leomo.com/v1
# Drop-in replacement — change the base URL only
curl https://api.leomo.com/v1/chat/completions \
  -H "Authorization: Bearer $LEOMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "leomo/b300-llama-405b",
    "messages": [
      {"role": "user", "content": "Summarise this report."}
    ],
    "stream": true
  }'
# Works with the official OpenAI SDK
from openai import OpenAI

client = OpenAI(
    base_url="https://api.leomo.com/v1",
    api_key="LEOMO_API_KEY"
)

resp = client.chat.completions.create(
    model="leomo/b300-llama-405b",
    messages=[{"role": "user",
               "content": "Summarise this report."}]
)
print(resp.choices[0].message.content)
Model Context Input / 1M Cached Input / 1M Output / 1M
leomo/b300-llama-405b Llama 3.1 405B · NVIDIA B300
128K $3.50 $0.70 $10.50
leomo/h200-deepseek-v3 DeepSeek-V3 · NVIDIA H200
128K $1.20 $0.24 $3.60
leomo/h100-qwen-72b Qwen 2.5 72B · NVIDIA H100
128K $0.60 $0.12 $1.80
leomo/mi300x-mixtral-8x22b Mixtral 8x22B · AMD MI300X
64K $0.90 $0.18 $2.70
leomo/gaudi3-embed-large Embedding model · Intel Gaudi 3
8K $0.05
* Pay-as-you-go pricing. Batch inference is billed at 50% of serverless rates. Provisioned throughput and private model hosting are quoted per deployment — volume discounts from 100M tokens/month. Request an API key
Three Ways to Work With Us

Pick the Level of Abstraction You Need

Own the hardware, own the model, or own nothing but the API call — same infrastructure underneath.

GPU Server Rental

Bare-metal NVIDIA, AMD and Intel servers by the hour or reserved monthly. Full root access, 400G InfiniBand, no hypervisor.

See rental plans

Operated AI Inference

Choose the model — open, commercial or custom. We run it on dedicated GPUs with autoscaling, observability and data residency controls built in.

Explore operated inference

Token API Platform

OpenAI-compatible endpoint billed per token. 100+ models, cached input discounts, batch inference at 50% off — no GPUs to manage.

View token pricing
Why LEOMO Company

Infrastructure Built for AI Teams

Latest Silicon in Stock

B300, H200, H100, MI300X and Gaudi 3 available now — not on a 6-month waitlist.

True Bare Metal

Physical, single-tenant servers with full root access. Your data never shares silicon.

Hardware, Model or API

Start on the Token API, move to operated inference or dedicated GPUs as volume grows — same team, same contract.

Engineers, Not Scripts

24/7 access to engineers who know CUDA, ROCm, vLLM and distributed training frameworks.

Solutions

Built for Your Workload

Optimised configurations matched to your application, from large language models to simulation.

AI Model Training

Train transformer, diffusion and multimodal models on clusters with massive HBM capacity and low-latency interconnect.

  • 8× GPU nodes with NVLink / Infinity Fabric
  • Up to 2.3TB HBM3e per node
  • 400G InfiniBand RDMA networking
  • PyTorch, TensorFlow and JAX ready
GPU per Training Node

Large-Scale Inference

Serve LLMs with massive context windows — on your own rented cluster, through operated inference, or via the Token API.

  • Up to 288GB HBM3e per GPU
  • FP4 / FP6 / FP8 quantisation support
  • Multi-node tensor parallelism
  • OpenAI-compatible API endpoints
288
GB Max HBM per GPU

Scientific Computing

Accelerate molecular dynamics, CFD and weather modelling with FP64-capable GPUs and ECC memory.

  • FP64 compute for simulation
  • ECC memory for data integrity
  • MPI and OpenMP support
  • Tuned for GROMACS, LAMMPS and VASP
FP64
Double Precision Ready

Film & 3D Rendering

GPU-accelerated rendering for VFX pipelines. Cut render times from hours to minutes with parallel GPU nodes.

  • Blender, Octane and Redshift support
  • RTX PRO 6000 and L40S GPUs
  • NVLink for multi-GPU rendering
  • On-demand burst rendering
10×
Faster Render Pipelines
How It Works

From Request to Running in 4 Steps

STEP 01

Share Your Workload

Tell us your model size, framework, expected throughput and timeline.

STEP 02

Pick Your Abstraction

Rent bare metal, hand us the model for operated inference, or take a Token API key.

STEP 03

Deploy & Validate

We provision the hardware, install drivers (CUDA / ROCm) and run benchmarks.

STEP 04

Scale Anytime

Add nodes, upgrade GPUs, or move your workload between rental, operated inference and Token API.

FAQ

Frequently Asked Questions

Everything you need to know about GPU rental, operated inference and the Token API.

All three run on the same LEOMO hardware — the difference is how much you manage:

  • GPU Server Rental — you get bare-metal servers with full root access. You install your own stack, run your own training or serving framework, and pay by the GPU-hour.
  • Operated AI Inference — you bring the model (including fine-tuned weights); we run the serving stack, autoscaling and monitoring on dedicated GPUs behind your own endpoint. Billed as reserved GPU capacity.
  • Token API Platform — you bring nothing but an API call. Shared pools, OpenAI-compatible endpoint, billed purely per input and output token.

Most customers start on the Token API for prototyping, then move to operated inference or dedicated rental once volume and control requirements justify it.

It depends on the configuration:

  • Token API: an API key is issued the same business day — no deployment needed.
  • On-demand rental: in-stock configurations are available in under 60 minutes.
  • Reserved / dedicated bare metal: typically provisioned within 24 hours of a signed order.
  • Operated inference: 3–7 business days for model onboarding, quantisation and load testing.
  • Multi-node clusters: 2–4 weeks, including fabric cabling and cluster-level validation.

Every hardware deployment includes driver installation, benchmark validation and a handover session with an engineer before you take control.

Our current rental fleet covers three vendors:

  • NVIDIA: B300, H200, H100 SXM, L40S, A100 SXM, RTX PRO 6000
  • AMD Instinct: MI300X (192GB HBM3)
  • Intel Gaudi: Gaudi 3 HL-338 (128GB HBM2e)

Yes — you can mix architectures inside one rented cluster. A common pattern is using B300 or H200 for large-model inference, H100 for training, L40S for lightweight inference and rendering, and MI300X for high-memory inference. We will confirm driver and framework compatibility for your stack before provisioning.

Yes. Operated AI Inference is designed for exactly that. You provide the weights (typically via a secure upload or a private object store bucket), and we handle:

  • Quantisation (FP8 / FP4 / INT4) and tensor parallelism configuration
  • Serving stack selection — vLLM, TensorRT-LLM or TGI depending on the architecture
  • Load testing and latency benchmarking before you go live
  • Rolling updates when you push a new checkpoint

Models run on dedicated, single-tenant GPUs and are never shared with other customers. Your weights are deleted on request to NIST 800-88 standards.

Every rental and operated inference deployment is single-tenant physical hardware — no hypervisor, no shared GPU, no noisy neighbours. On the Token API, your prompts and completions are never used to train models and are not retained beyond the request unless you explicitly enable logging. Security features include:

  • Full-disk encryption and secure erase on decommission
  • Private VLANs, VPN gateways and firewall rules you control
  • Optional IPMI / BMC access behind a private management network
  • Scoped API keys with per-key rate limits and usage caps
  • NDA and Data Processing Agreements available before any data is uploaded

We never inspect or copy the contents of your storage, and data is wiped to NIST 800-88 standards when a deployment ends.

The Token API is billed purely on input and output tokens at the rates published in the model catalog — there is no hourly GPU charge and no minimum commit. Cached input tokens are billed at a discount when the same prompt prefix repeats. Batch inference is billed at 50% of serverless rates. Usage is metered per request and visible in your dashboard in near real time; you can set hard spend caps per API key.

All plans, including pay-as-you-go API access, include 24/7 access to human engineers via ticket and email, plus a 99.9% API availability SLA. Hardware faults are replaced within 4 hours on reserved, operated inference and enterprise cluster deployments.

Optional managed services cover cluster deployment, CUDA / ROCm driver tuning, distributed training setup (NCCL, RCCL, DeepSpeed), private model hosting, performance benchmarking and migration assistance. Volume and reserved customers also receive a named account engineer and quarterly capacity reviews.

Rent the Hardware, or Just Call the API

Tell us about your workload: model size, framework, throughput and timeline. We will recommend the right abstraction — bare-metal rental, operated inference, or a Token API key — within one business day.