B300
Frontier training and large-scale inference with 288GB HBM3e and 2× Hopper FP4 throughput.
- Memory288GB HBM3e
- Bandwidth8.2 TB/s
- FP42× Hopper
- ConfigHGX / SXM
LEOMO Company operates NVIDIA, AMD and Intel bare-metal infrastructure. Rent the hardware by the hour, or call our OpenAI-compatible Token API and pay only for what you generate.
Bare-metal NVIDIA, AMD and Intel servers by the hour or reserved monthly. Full root access, no hypervisor.
See rental plansChoose your model. We run it on dedicated GPUs with autoscaling, observability and data residency controls built in.
Explore operated inferenceOpenAI-compatible endpoint, billed per token. No GPUs to manage, no idle capacity to pay for.
View token pricingFilter by manufacturer to compare specifications, on-demand hourly rates and reserved monthly rates. Pricing reflects current US and Japan cloud market benchmarks.
Frontier training and large-scale inference with 288GB HBM3e and 2× Hopper FP4 throughput.
141GB HBM3e delivers up to 1.9× faster LLM inference than H100 at competitive cost.
The industry-standard AI training accelerator, supported by every major framework and cloud toolchain.
48GB GDDR6 with strong FP16/INT8 throughput — the best price-performance entry point for AI inference.
80GB HBM2e with proven reliability — still the workhorse for fine-tuning and mid-scale training.
96GB GDDR7 workstation GPU — the most cost-effective entry point for AI, rendering and visualisation.
No capital expenditure, no hardware to maintain. Choose on-demand flexibility or lock in reserved capacity at up to 40% below list rate.
Hourly billing, deploy in minutes, terminate anytime. Ideal for experiments and burst workloads.
Commit for 1–12 months and save up to 40%. For steady production inference and training pipelines.
Multi-node deployments with NVLink, Infinity Fabric and 400G InfiniBand, validated for distributed training.
Deploy any open, commercial or custom model on managed inference infrastructure. Autoscaling, observability and data residency controls are built in — you never touch a GPU.
Run an open-source model, a commercial model, or your own fine-tuned weights through a managed endpoint. LEOMO operates the serving environment, handles autoscaling, and keeps your application interface stable while models and infrastructure evolve underneath.
Call 100+ models through a shared per-token API. No provisioning, no replicas to size, no minimum cost — and no idle GPUs to pay for.
Change one base URL and your existing OpenAI SDK code keeps working. The Token API is the fastest way to run inference on LEOMO hardware — you are billed on input and output tokens, with no capacity to reserve and no commitment.
https://api.leomo.com/v1.
# Drop-in replacement — change the base URL only curl https://api.leomo.com/v1/chat/completions \ -H "Authorization: Bearer $LEOMO_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "leomo/b300-llama-405b", "messages": [ {"role": "user", "content": "Summarise this report."} ], "stream": true }'
# Works with the official OpenAI SDK from openai import OpenAI client = OpenAI( base_url="https://api.leomo.com/v1", api_key="LEOMO_API_KEY" ) resp = client.chat.completions.create( model="leomo/b300-llama-405b", messages=[{"role": "user", "content": "Summarise this report."}] ) print(resp.choices[0].message.content)
| Model | Context | Input / 1M | Cached Input / 1M | Output / 1M |
|---|---|---|---|---|
|
leomo/b300-llama-405b
Llama 3.1 405B · NVIDIA B300
|
128K | $3.50 | $0.70 | $10.50 |
|
leomo/h200-deepseek-v3
DeepSeek-V3 · NVIDIA H200
|
128K | $1.20 | $0.24 | $3.60 |
|
leomo/h100-qwen-72b
Qwen 2.5 72B · NVIDIA H100
|
128K | $0.60 | $0.12 | $1.80 |
|
leomo/mi300x-mixtral-8x22b
Mixtral 8x22B · AMD MI300X
|
64K | $0.90 | $0.18 | $2.70 |
|
leomo/gaudi3-embed-large
Embedding model · Intel Gaudi 3
|
8K | $0.05 | — | — |
Own the hardware, own the model, or own nothing but the API call — same infrastructure underneath.
Bare-metal NVIDIA, AMD and Intel servers by the hour or reserved monthly. Full root access, 400G InfiniBand, no hypervisor.
See rental plansChoose the model — open, commercial or custom. We run it on dedicated GPUs with autoscaling, observability and data residency controls built in.
Explore operated inferenceOpenAI-compatible endpoint billed per token. 100+ models, cached input discounts, batch inference at 50% off — no GPUs to manage.
View token pricingB300, H200, H100, MI300X and Gaudi 3 available now — not on a 6-month waitlist.
Physical, single-tenant servers with full root access. Your data never shares silicon.
Start on the Token API, move to operated inference or dedicated GPUs as volume grows — same team, same contract.
24/7 access to engineers who know CUDA, ROCm, vLLM and distributed training frameworks.
Optimised configurations matched to your application, from large language models to simulation.
Train transformer, diffusion and multimodal models on clusters with massive HBM capacity and low-latency interconnect.
Serve LLMs with massive context windows — on your own rented cluster, through operated inference, or via the Token API.
Accelerate molecular dynamics, CFD and weather modelling with FP64-capable GPUs and ECC memory.
GPU-accelerated rendering for VFX pipelines. Cut render times from hours to minutes with parallel GPU nodes.
Tell us your model size, framework, expected throughput and timeline.
Rent bare metal, hand us the model for operated inference, or take a Token API key.
We provision the hardware, install drivers (CUDA / ROCm) and run benchmarks.
Add nodes, upgrade GPUs, or move your workload between rental, operated inference and Token API.
Everything you need to know about GPU rental, operated inference and the Token API.
All three run on the same LEOMO hardware — the difference is how much you manage:
Most customers start on the Token API for prototyping, then move to operated inference or dedicated rental once volume and control requirements justify it.
It depends on the configuration:
Every hardware deployment includes driver installation, benchmark validation and a handover session with an engineer before you take control.
Our current rental fleet covers three vendors:
Yes — you can mix architectures inside one rented cluster. A common pattern is using B300 or H200 for large-model inference, H100 for training, L40S for lightweight inference and rendering, and MI300X for high-memory inference. We will confirm driver and framework compatibility for your stack before provisioning.
Yes. Operated AI Inference is designed for exactly that. You provide the weights (typically via a secure upload or a private object store bucket), and we handle:
Models run on dedicated, single-tenant GPUs and are never shared with other customers. Your weights are deleted on request to NIST 800-88 standards.
Every rental and operated inference deployment is single-tenant physical hardware — no hypervisor, no shared GPU, no noisy neighbours. On the Token API, your prompts and completions are never used to train models and are not retained beyond the request unless you explicitly enable logging. Security features include:
We never inspect or copy the contents of your storage, and data is wiped to NIST 800-88 standards when a deployment ends.
The Token API is billed purely on input and output tokens at the rates published in the model catalog — there is no hourly GPU charge and no minimum commit. Cached input tokens are billed at a discount when the same prompt prefix repeats. Batch inference is billed at 50% of serverless rates. Usage is metered per request and visible in your dashboard in near real time; you can set hard spend caps per API key.
All plans, including pay-as-you-go API access, include 24/7 access to human engineers via ticket and email, plus a 99.9% API availability SLA. Hardware faults are replaced within 4 hours on reserved, operated inference and enterprise cluster deployments.
Optional managed services cover cluster deployment, CUDA / ROCm driver tuning, distributed training setup (NCCL, RCCL, DeepSpeed), private model hosting, performance benchmarking and migration assistance. Volume and reserved customers also receive a named account engineer and quarterly capacity reviews.