AI InferenceComing soon
Model endpoints with autoscaling and request metering.
Managed inference endpoints on top of GPU Cloud: an HTTPS endpoint, autoscaling by concurrency, request-level metering and logs — without operating a serving stack yourself.
At a glance
Coming soon
- Management
- platform
- Status
- Coming soon
Included
- Managed endpoints
- Autoscaling
- Request metering
- Logs and metrics
Why this
What AI Inference gives you
Endpoint, not a cluster
Deploy a model and get a URL. Scaling, health and restarts are the platform's job.
Observable by request
Latency, throughput and cost are attributed per request rather than per node.
Not yet available
AI Inference is in build
We do not sell capacity before it exists. Leave your address and we will contact you the moment this opens, with launch pricing.
Best for
Who runs this
- Production model serving
- Spiky inference traffic
- Teams without an MLOps function
Included
What comes with it
- Managed endpoints
- Autoscaling
- Request metering
- Logs and metrics
Questions
Frequently asked
Anything you can containerise. We publish supported serving runtimes when the product opens.
Build it on ScodyX Cloud
Create an account, configure what you need and watch it provision. No sales call required for anything with a published price.