Scody Cloud
AI InferenceComing soon

Model endpoints with autoscaling and request metering.

Managed inference endpoints on top of GPU Cloud: an HTTPS endpoint, autoscaling by concurrency, request-level metering and logs — without operating a serving stack yourself.

At a glance

Coming soon

Management
platform
Status
Coming soon

Included

  • Managed endpoints
  • Autoscaling
  • Request metering
  • Logs and metrics
Why this

What AI Inference gives you

Endpoint, not a cluster

Deploy a model and get a URL. Scaling, health and restarts are the platform's job.

Observable by request

Latency, throughput and cost are attributed per request rather than per node.

Not yet available

AI Inference is in build

We do not sell capacity before it exists. Leave your address and we will contact you the moment this opens, with launch pricing.

Best for

Who runs this

  • Production model serving
  • Spiky inference traffic
  • Teams without an MLOps function
Included

What comes with it

  • Managed endpoints
  • Autoscaling
  • Request metering
  • Logs and metrics
Questions

Frequently asked

Anything you can containerise. We publish supported serving runtimes when the product opens.

Build it on ScodyX Cloud

Create an account, configure what you need and watch it provision. No sales call required for anything with a published price.