# GPU rental reference

A rental is two objects:

- A **workload** describes *what* to run: the image, the hardware, SSH keys, ports and disk. You can reuse it.
- A **deployment** is one running instance of a workload. It is billed per minute while it is `ready`. Pulling the image and starting up are free, and so is a deployment that fails.

All endpoints live under `https://api.green-compute.com` and take `Authorization: Bearer <key>`, except where marked public.

## Hardware

| GPU id | VRAM | Price per GPU-hour |
|---|---|---|
| `rtx4090` | 24 GB | $0.40 |
| `rtx5090` | 32 GB | $0.70 |

These are the current figures. The authoritative live source is `GET /platform/pricing` (public), which is what billing actually uses. To see how many GPUs are free right now, call `GET https://control.green-compute.com/platform/v1/gpu-pool` (public).

GPU ids are matched without regard to case or separators, so `rtx4090`, `RTX 4090` and `rtx-4090` are the same. `GET /platform/nodes/supported` (public) lists the ids you can rent.

> **RTX 5090 needs a CUDA 12.8+ image.** The 5090 is Blackwell (`sm_120`). Images built for older CUDA, such as `pytorch/pytorch:2.2.0-cuda12.1-*`, start fine but fail the moment you use the GPU. `pytorch/pytorch:2.7.0-cuda12.8-cudnn9-runtime` works on both cards.

## Create a workload — `POST /platform/workloads`

```json
{
  "name": "my-gpu-box",
  "kind": "pod",
  "image": "pytorch/pytorch:2.7.0-cuda12.8-cudnn9-runtime",
  "requirements": {
    "gpu_count": 1,
    "supported_gpu_models": ["rtx4090"],
    "min_vram_gb_per_gpu": 24,
    "cpu_cores": 8,
    "memory_gb": 32
  },
  "metadata": {
    "ssh_public_keys": ["ssh-ed25519 AAAA... you@laptop"],
    "requested_ports": [8888],
    "volume_size_gb": 50
  }
}
```

| Field | Meaning |
|---|---|
| `kind` | Must be `"pod"` for a rental. |
| `image` | Any public Docker image. SSH is injected at start-up, so the image does not need its own SSH server. |
| `requirements.gpu_count` | 1 to 8 GPUs on a single machine. |
| `requirements.supported_gpu_models` | GPU ids you accept. Leave it empty to accept any; the quoted rate is then the highest. |
| `requirements.cpu_cores`, `memory_gb` | Capped at the machine's fair share for the number of GPUs you take. |
| `metadata.ssh_public_keys` | Your public keys. They are authorised for `root`. Recommended: no private key ever has to leave your machine. |
| `metadata.requested_ports` | Up to 10 container ports to expose (for example Jupyter on 8888). Their public host ports appear in the deployment's `port_mappings`. |
| `metadata.volume_size_gb` | Disk for `/workspace`, 10 to 2000 GB. Default 50. |

The response contains `workload_id`.

## Start it — `POST /platform/deployments`

```json
{ "workload_id": "…", "save_on_exhaustion": true }
```

The response is the deployment. Fields worth reading:

| Field | Meaning |
|---|---|
| `deployment_id` | Use this for every later call. |
| `state` | See the lifecycle below. |
| `hourly_rate_cents` | Price per GPU-hour, in cents. At creation it is the most this deployment can cost; once a GPU is assigned, it becomes that card's exact rate, which is never higher. |
| `deployment_fee_usd` | Always `0`. There is no deployment fee. |
| `port_mappings` | `{container_port: public_host_port}` once running. |
| `last_error` | Set if placement or start-up failed. |

Starting requires at least one hour of the quoted rate times `gpu_count` in your balance. Otherwise you get `402` (see [Errors](https://www.green-compute.com/docs/errors.md)).

`save_on_exhaustion` (default `true`): if your balance runs out, the pod is suspended with its disk kept, rather than deleted, so you can resume it after topping up.

## Lifecycle

```
pending → scheduled → pulling → starting → ready
                                             │
                     suspended ◄── (balance ran out, save_on_exhaustion)
                                             │
                                  terminated ◄── DELETE
```

`failed` means it could not be placed or started. Read `last_error`. Poll `GET /platform/deployments/{id}` every 5–10 seconds; most rentals are `ready` within one to three minutes.

## Connect — `GET /platform/deployments/{id}/ssh`

```json
{
  "ssh_host": "…",
  "ssh_port": 40123,
  "ssh_username": "root",
  "ssh_command": "ssh root@… -p 40123",
  "private_key": "-----BEGIN OPENSSH PRIVATE KEY-----…"
}
```

If you supplied `ssh_public_keys`, connect with your own key and ignore `private_key`. The platform-generated key is returned **only** by this endpoint, and every retrieval is audit-logged. It never appears in create, get, list or delete responses.

Inside the pod:
- the image's environment is preserved, so `python`, `conda` and `CUDA_HOME` behave as the image defines them, including for one-off commands like `ssh … python train.py`;
- `/workspace` is your persistent disk;
- `nvidia-smi` shows only the GPUs you rented.

## Other calls

| Call | Does |
|---|---|
| `GET /platform/deployments` | Lists your deployments. |
| `GET /platform/deployments/{id}` | One deployment. |
| `DELETE /platform/deployments/{id}` | Stops it and **stops billing**. |
| `POST /platform/deployments/{id}/resume` | Restarts a `suspended` pod in place, with its disk intact. Needs one hour of balance. |
| `DELETE /platform/workloads/{id}` | Deletes the workload definition. |

## Limits

- 30 workload creates and 30 deployment creates per minute per API key.
- One GPU machine per deployment; `gpu_count` up to the machine size (8).
