# Green Compute documentation > Generated from https://www.green-compute.com/docs. Index: https://www.green-compute.com/llms.txt # Quickstart: rent a GPU with the API This page takes you from an API key to a running RTX 4090 you can SSH into, then shuts it down. It uses plain HTTP, so it works the same from `curl`, Python, or an AI agent. Allow about five minutes. - **API base URL:** `https://api.green-compute.com` - **Auth:** `Authorization: Bearer ` on every request except the public ones marked below ## 0. One-time setup (a human does this once) 1. Sign in at [green-compute.com/login](https://www.green-compute.com/login). 2. Add credit on the [Billing](https://www.green-compute.com/billing) page (card or TAO). Starting a rental needs at least **one hour** of its price in your balance. 3. Create an API key under [Settings → API keys](https://www.green-compute.com/settings). Everything after this step is API-only. ```bash export GC_KEY="paste-your-api-key" export GC_API="https://api.green-compute.com" ``` ## 1. Check price and availability (public, no key) ```bash curl -s "$GC_API/platform/pricing" curl -s https://control.green-compute.com/platform/v1/gpu-pool ``` `/platform/pricing` is the exact rate billing charges. `gpu-pool` shows how many GPUs of each model are free right now. GPU ids are `rtx4090` and `rtx5090`, and any spelling works on input (`"RTX 4090"`, `"rtx-4090"`). ## 2. Describe the machine you want A *workload* says what to run: the image and the hardware. Put your own SSH **public** key in `metadata.ssh_public_keys` so you can log in with a key you already hold. ```bash WORKLOAD_ID=$(curl -s -X POST "$GC_API/platform/workloads" \ -H "Authorization: Bearer $GC_KEY" -H "Content-Type: application/json" \ -d '{ "name": "my-gpu-box", "kind": "pod", "image": "pytorch/pytorch:2.7.0-cuda12.8-cudnn9-runtime", "requirements": { "gpu_count": 1, "supported_gpu_models": ["rtx4090"], "min_vram_gb_per_gpu": 24, "cpu_cores": 8, "memory_gb": 32 }, "metadata": { "ssh_public_keys": ["'"$(cat ~/.ssh/id_ed25519.pub)"'"], "volume_size_gb": 50 } }' | python3 -c 'import sys,json; print(json.load(sys.stdin)["workload_id"])') echo "$WORKLOAD_ID" ``` ## 3. Start it A *deployment* is a running instance of the workload. Billing starts when it reaches `ready`, so time spent pulling the image is free. ```bash DEPLOYMENT_ID=$(curl -s -X POST "$GC_API/platform/deployments" \ -H "Authorization: Bearer $GC_KEY" -H "Content-Type: application/json" \ -d "{\"workload_id\": \"$WORKLOAD_ID\"}" \ | python3 -c 'import sys,json; d=json.load(sys.stdin); print(d["deployment_id"]); print("rate cents/GPU/hr:", d["hourly_rate_cents"], file=sys.stderr)') ``` `hourly_rate_cents` in the response is the most this deployment can cost per GPU-hour. With a single GPU model it is the exact rate. Once a GPU is assigned, the field updates to that card's rate, which is never higher. There is no deployment fee: `deployment_fee_usd` is always `0`. ## 4. Wait until it is ready ```bash while true; do STATE=$(curl -s "$GC_API/platform/deployments/$DEPLOYMENT_ID" -H "Authorization: Bearer $GC_KEY" \ | python3 -c 'import sys,json; print(json.load(sys.stdin)["state"])') echo "$STATE" case "$STATE" in ready) break ;; failed|terminated) exit 1 ;; esac sleep 10 done ``` States run `pending → scheduled → pulling → starting → ready`. Usually this takes one to three minutes; the first pull of a large image takes longer. ## 5. Connect ```bash curl -s "$GC_API/platform/deployments/$DEPLOYMENT_ID/ssh" -H "Authorization: Bearer $GC_KEY" # {"ssh_host": "...", "ssh_port": 40123, "ssh_username": "root", "ssh_command": "ssh root@... -p 40123", "private_key": "..."} ssh root@ -p nvidia-smi ``` Your public key from step 2 is already authorised, so you don't need the `private_key` field. It is a platform-generated key, returned only by this endpoint, for clients that didn't supply their own. The image's environment carries over into SSH sessions, so `python` and `conda` work exactly as the image defines them. Persistent storage is mounted at `/workspace`. ## 6. Stop it (this stops billing) ```bash curl -s -X DELETE "$GC_API/platform/deployments/$DEPLOYMENT_ID" -H "Authorization: Bearer $GC_KEY" ``` Billing is per minute, so you pay only for the time it ran. ## Next - [GPU rental reference](https://www.green-compute.com/docs/gpu-rental.md): every field, ports, disk, resume, 5090 notes - [Errors](https://www.green-compute.com/docs/errors.md): what each status code means and how to recover - [Inference API](https://www.green-compute.com/docs/inference.md): OpenAI-compatible chat completions --- # For AI agents This page is for an AI agent (or its developer) that has been pointed at Green Compute and needs to use it without a human clicking through a dashboard. ## What Green Compute is GPU compute on Bittensor subnet 110, run on renewable and biogas power. Two products, both over one HTTP API: 1. **GPU rental:** an SSH-accessible container on an RTX 4090 or 5090, billed per minute. See [Quickstart](https://www.green-compute.com/docs/quickstart.md). 2. **Inference:** OpenAI-compatible chat completions. See [Inference](https://www.green-compute.com/docs/inference.md). ## Machine-readable entry points | What | URL | Auth | |---|---|---| | This documentation, as one file | `https://www.green-compute.com/llms-full.txt` | none | | Documentation index | `https://www.green-compute.com/llms.txt` | none | | Any docs page as Markdown | append `.md`, e.g. `https://www.green-compute.com/docs/quickstart.md` | none | | OpenAPI spec | `https://api.green-compute.com/openapi.json` | none | | API index | `https://api.green-compute.com/` | none | | Prices | `https://api.green-compute.com/platform/pricing` | none | | Rentable GPU ids | `https://api.green-compute.com/platform/nodes/supported` | none | | Free GPUs right now | `https://control.green-compute.com/platform/v1/gpu-pool` | none | ## What needs a human, once Creating the account, adding credit and creating the API key happen in the browser. After that, everything is API. If you have no key, ask your user for one. Point them at [green-compute.com/settings](https://www.green-compute.com/settings) and [green-compute.com/billing](https://www.green-compute.com/billing). ## Rules of thumb - **Use your own SSH public key** in `metadata.ssh_public_keys`. Then you never need to fetch, store or log a private key. - **Budget from the API, not from prose.** `hourly_rate_cents` on a new deployment is the most it can cost per GPU-hour; `/platform/pricing` is what billing charges. - **Always clean up.** `DELETE /platform/deployments/{id}` stops billing. Wrap your work so this runs even on failure. - **Poll, don't sleep blindly.** `GET /platform/deployments/{id}` every 5–10 s until `state` is `ready`; stop on `failed` and read `last_error`. - **Errors are actionable.** A `400` for an unavailable GPU lists the GPUs that are available; a `402` gives the exact amount to top up. See [Errors](https://www.green-compute.com/docs/errors.md). - **For an RTX 5090, use a CUDA 12.8+ image.** `pytorch/pytorch:2.7.0-cuda12.8-cudnn9-runtime` works on both cards. - **One-off commands work over SSH.** `ssh root@host -p port 'python train.py'` uses the image's `PATH`, so conda/PyTorch images behave as documented. ## Minimal loop (Python, standard library only) ```python import json, os, time, urllib.request API = "https://api.green-compute.com" KEY = os.environ["GC_KEY"] def call(method, path, body=None): req = urllib.request.Request( API + path, method=method, data=json.dumps(body).encode() if body is not None else None, headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}, ) with urllib.request.urlopen(req) as r: return json.loads(r.read() or b"null") pubkey = open(os.path.expanduser("~/.ssh/id_ed25519.pub")).read().strip() wl = call("POST", "/platform/workloads", { "name": "agent-job", "kind": "pod", "image": "pytorch/pytorch:2.7.0-cuda12.8-cudnn9-runtime", "requirements": {"gpu_count": 1, "supported_gpu_models": ["rtx4090"], "min_vram_gb_per_gpu": 24, "cpu_cores": 8, "memory_gb": 32}, "metadata": {"ssh_public_keys": [pubkey]}, }) dep = call("POST", "/platform/deployments", {"workload_id": wl["workload_id"]}) try: while (d := call("GET", f"/platform/deployments/{dep['deployment_id']}"))["state"] not in ("ready", "failed"): time.sleep(10) if d["state"] == "failed": raise RuntimeError(d.get("last_error")) ssh = call("GET", f"/platform/deployments/{dep['deployment_id']}/ssh") print("connect with:", ssh["ssh_command"]) # ... do the work over SSH ... finally: call("DELETE", f"/platform/deployments/{dep['deployment_id']}") ``` --- # GPU rental reference A rental is two objects: - A **workload** describes *what* to run: the image, the hardware, SSH keys, ports and disk. You can reuse it. - A **deployment** is one running instance of a workload. It is billed per minute while it is `ready`. Pulling the image and starting up are free, and so is a deployment that fails. All endpoints live under `https://api.green-compute.com` and take `Authorization: Bearer `, except where marked public. ## Hardware | GPU id | VRAM | Price per GPU-hour | |---|---|---| | `rtx4090` | 24 GB | $0.40 | | `rtx5090` | 32 GB | $0.70 | These are the current figures. The authoritative live source is `GET /platform/pricing` (public), which is what billing actually uses. To see how many GPUs are free right now, call `GET https://control.green-compute.com/platform/v1/gpu-pool` (public). GPU ids are matched without regard to case or separators, so `rtx4090`, `RTX 4090` and `rtx-4090` are the same. `GET /platform/nodes/supported` (public) lists the ids you can rent. > **RTX 5090 needs a CUDA 12.8+ image.** The 5090 is Blackwell (`sm_120`). Images built for older CUDA, such as `pytorch/pytorch:2.2.0-cuda12.1-*`, start fine but fail the moment you use the GPU. `pytorch/pytorch:2.7.0-cuda12.8-cudnn9-runtime` works on both cards. ## Create a workload — `POST /platform/workloads` ```json { "name": "my-gpu-box", "kind": "pod", "image": "pytorch/pytorch:2.7.0-cuda12.8-cudnn9-runtime", "requirements": { "gpu_count": 1, "supported_gpu_models": ["rtx4090"], "min_vram_gb_per_gpu": 24, "cpu_cores": 8, "memory_gb": 32 }, "metadata": { "ssh_public_keys": ["ssh-ed25519 AAAA... you@laptop"], "requested_ports": [8888], "volume_size_gb": 50 } } ``` | Field | Meaning | |---|---| | `kind` | Must be `"pod"` for a rental. | | `image` | Any public Docker image. SSH is injected at start-up, so the image does not need its own SSH server. | | `requirements.gpu_count` | 1 to 8 GPUs on a single machine. | | `requirements.supported_gpu_models` | GPU ids you accept. Leave it empty to accept any; the quoted rate is then the highest. | | `requirements.cpu_cores`, `memory_gb` | Capped at the machine's fair share for the number of GPUs you take. | | `metadata.ssh_public_keys` | Your public keys. They are authorised for `root`. Recommended: no private key ever has to leave your machine. | | `metadata.requested_ports` | Up to 10 container ports to expose (for example Jupyter on 8888). Their public host ports appear in the deployment's `port_mappings`. | | `metadata.volume_size_gb` | Disk for `/workspace`, 10 to 2000 GB. Default 50. | The response contains `workload_id`. ## Start it — `POST /platform/deployments` ```json { "workload_id": "…", "save_on_exhaustion": true } ``` The response is the deployment. Fields worth reading: | Field | Meaning | |---|---| | `deployment_id` | Use this for every later call. | | `state` | See the lifecycle below. | | `hourly_rate_cents` | Price per GPU-hour, in cents. At creation it is the most this deployment can cost; once a GPU is assigned, it becomes that card's exact rate, which is never higher. | | `deployment_fee_usd` | Always `0`. There is no deployment fee. | | `port_mappings` | `{container_port: public_host_port}` once running. | | `last_error` | Set if placement or start-up failed. | Starting requires at least one hour of the quoted rate times `gpu_count` in your balance. Otherwise you get `402` (see [Errors](https://www.green-compute.com/docs/errors.md)). `save_on_exhaustion` (default `true`): if your balance runs out, the pod is suspended with its disk kept, rather than deleted, so you can resume it after topping up. ## Lifecycle ``` pending → scheduled → pulling → starting → ready │ suspended ◄── (balance ran out, save_on_exhaustion) │ terminated ◄── DELETE ``` `failed` means it could not be placed or started. Read `last_error`. Poll `GET /platform/deployments/{id}` every 5–10 seconds; most rentals are `ready` within one to three minutes. ## Connect — `GET /platform/deployments/{id}/ssh` ```json { "ssh_host": "…", "ssh_port": 40123, "ssh_username": "root", "ssh_command": "ssh root@… -p 40123", "private_key": "-----BEGIN OPENSSH PRIVATE KEY-----…" } ``` If you supplied `ssh_public_keys`, connect with your own key and ignore `private_key`. The platform-generated key is returned **only** by this endpoint, and every retrieval is audit-logged. It never appears in create, get, list or delete responses. Inside the pod: - the image's environment is preserved, so `python`, `conda` and `CUDA_HOME` behave as the image defines them, including for one-off commands like `ssh … python train.py`; - `/workspace` is your persistent disk; - `nvidia-smi` shows only the GPUs you rented. ## Other calls | Call | Does | |---|---| | `GET /platform/deployments` | Lists your deployments. | | `GET /platform/deployments/{id}` | One deployment. | | `DELETE /platform/deployments/{id}` | Stops it and **stops billing**. | | `POST /platform/deployments/{id}/resume` | Restarts a `suspended` pod in place, with its disk intact. Needs one hour of balance. | | `DELETE /platform/workloads/{id}` | Deletes the workload definition. | ## Limits - 30 workload creates and 30 deployment creates per minute per API key. - One GPU machine per deployment; `gpu_count` up to the machine size (8). --- # Inference API (OpenAI-compatible) Green Compute serves open models behind an OpenAI-compatible API. Any OpenAI SDK or tool works once you change the base URL and key. - **Base URL:** `https://api.green-compute.com/v1` - **Auth:** `Authorization: Bearer ` ## List models ```bash curl -s https://api.green-compute.com/v1/models -H "Authorization: Bearer $GC_KEY" ``` Use an `id` from this list verbatim as `model`. Per-token prices for each listed model are in `GET https://api.green-compute.com/platform/pricing` (public). ## Chat completions ```bash curl -s https://api.green-compute.com/v1/chat/completions \ -H "Authorization: Bearer $GC_KEY" -H "Content-Type: application/json" \ -d '{ "model": "", "messages": [{"role": "user", "content": "Say hello in five words."}], "max_tokens": 200 }' ``` ### Python ```python from openai import OpenAI client = OpenAI(base_url="https://api.green-compute.com/v1", api_key="YOUR_KEY") resp = client.chat.completions.create( model="", messages=[{"role": "user", "content": "Say hello in five words."}], max_tokens=200, ) print(resp.choices[0].message.content) ``` ### Streaming Set `"stream": true`. Responses are standard server-sent events, ending with `data: [DONE]`. The final chunk carries `usage`. If you disconnect mid-stream, you are billed only for the tokens actually generated. ### Tool calling Standard OpenAI `tools` / `tool_choice` are passed through. `tool_calls` come back in the usual place on the message. ### Reasoning models Some models think before answering. For those: - the thinking arrives in `message.reasoning_content` (not `reasoning`) and the answer in `message.content`; - **thinking counts against `max_tokens`**, so budget generously (2000+). If `content` comes back empty, the model ran out of budget while thinking; raise `max_tokens` and retry. ## Using it from coding tools Any tool that accepts a custom OpenAI base URL works: | Setting | Value | |---|---| | Base URL | `https://api.green-compute.com/v1` | | API key | your Green Compute key | | Model | an `id` from `/v1/models` | ## Billing Per token, input and output priced separately, with a small minimum charge per request. See `GET /platform/pricing` for the live figures. --- # Pricing ## GPU rental | GPU | VRAM | Per GPU-hour | |---|---|---| | RTX 4090 (`rtx4090`) | 24 GB | **$0.40** | | RTX 5090 (`rtx5090`) | 32 GB | **$0.70** | - **Billed per minute** while a deployment is `ready`. Image pulls, start-up and failed deployments are free. - **No deployment fee, no egress fee, no minimum term.** Stop with `DELETE /platform/deployments/{id}` and billing stops. - To start a rental, your balance must cover **one hour** of its price. - The price is fixed when a GPU is assigned to your deployment, so later list-price changes don't affect a running rental. ## Inference Per token, input and output priced separately, with a minimum charge of $0.01 per request. Prices vary by model. ## Live, machine-readable prices The figures above can change. The authoritative source, which billing itself uses, is public and needs no key: ```bash curl -s https://api.green-compute.com/platform/pricing ``` ```json { "currency": "USD", "gpu_rental": { "billing": "per minute while running, per GPU", "deployment_fee_usd": 0, "gpus": [ {"gpu_model": "rtx4090", "vram_gb": 24, "cents_per_gpu_hour": 40, "usd_per_gpu_hour": 0.4}, {"gpu_model": "rtx5090", "vram_gb": 32, "cents_per_gpu_hour": 70, "usd_per_gpu_hour": 0.7} ] }, "inference": { "billing": "per token; minimum charge per request", "minimum_charge_usd": 0.01, "models": [{"model": "…", "usd_per_million_input_tokens": 0.2, "usd_per_million_output_tokens": 0.6}] } } ``` ## Paying Top up on the [Billing](https://www.green-compute.com/billing) page with a card or TAO. Your balance is drawn down per minute by rentals and per request by inference. --- # Errors and how to recover Every error is JSON with a `detail` field. Where there is something to act on, `detail` is an object with the specifics. | Status | When | What to do | |---|---|---| | `400` | A request is invalid, for example no requested GPU model exists in the fleet. | Read `detail`. For GPUs, `detail.available` lists the ids you can use. | | `401` | The API key is missing or wrong. | Send `Authorization: Bearer `. Create a key under [Settings](https://www.green-compute.com/settings). | | `402` | Your balance is below one hour of the rental's price. | Top up on [Billing](https://www.green-compute.com/billing). `detail` gives the exact amounts. | | `403` | Your key can't access that resource. | Use a key from the account that owns it. | | `404` | The deployment or workload doesn't exist, or isn't yours. On `/ssh`, the pod isn't ready yet. | Check the id; for `/ssh`, wait for `state: "ready"`. | | `409` | The action conflicts with the current state, for example resuming a pod whose machine is full. | Read `detail`; for resume, try again later or start a new rental. | | `422` | The request body doesn't match the schema. | `detail` lists each bad field and why. | | `429` | Rate limit (30 creates per minute per key). | Wait and retry with back-off. | ## The two you will see most ### `400` — GPU not available ```json { "detail": { "message": "none of the requested GPU models are available to rent", "requested": ["h100"], "available": ["rtx4090", "rtx5090"], "hint": "set supported_gpu_models to one of `available` (any spelling); see GET /platform/pricing" } } ``` This is returned immediately when you start the deployment, so you never wait on a rental that could never be placed. ### `402` — not enough balance ```json { "detail": { "message": "insufficient balance to start this rental", "required_cents": 40, "current_cents": 12, "rate_cents_per_hour": 40, "gpu_count": 1, "requested_instances": 1 } } ``` You need `required_cents - current_cents` more, here $0.28. Top up, then retry the same request. ## A deployment that ends up `failed` Read `last_error` on `GET /platform/deployments/{id}`. The most common causes: - **The image doesn't exist or is private.** Check the image name; only public images can be pulled. - **The image's CUDA is too old for an RTX 5090.** Use a CUDA 12.8+ image (see [GPU rental](https://www.green-compute.com/docs/gpu-rental.md)). A failed deployment is not billed.