Inference API (OpenAI-compatible)
Green Compute serves open models behind an OpenAI-compatible API. Any OpenAI SDK or tool works once you change the base URL and key.
- Base URL:
https://api.green-compute.com/v1 - Auth:
Authorization: Bearer <your API key>
List models
curl -s https://api.green-compute.com/v1/models -H "Authorization: Bearer $GC_KEY"
Use an id from this list verbatim as model. Per-token prices for each listed model are in GET https://api.green-compute.com/platform/pricing (public).
Chat completions
curl -s https://api.green-compute.com/v1/chat/completions \
-H "Authorization: Bearer $GC_KEY" -H "Content-Type: application/json" \
-d '{
"model": "<id from /v1/models>",
"messages": [{"role": "user", "content": "Say hello in five words."}],
"max_tokens": 200
}'
Python
from openai import OpenAI
client = OpenAI(base_url="https://api.green-compute.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="<id from /v1/models>",
messages=[{"role": "user", "content": "Say hello in five words."}],
max_tokens=200,
)
print(resp.choices[0].message.content)
Streaming
Set "stream": true. Responses are standard server-sent events, ending with data: [DONE]. The final chunk carries usage. If you disconnect mid-stream, you are billed only for the tokens actually generated.
Tool calling
Standard OpenAI tools / tool_choice are passed through. tool_calls come back in the usual place on the message.
Reasoning models
Some models think before answering. For those:
- the thinking arrives in
message.reasoning_content(notreasoning) and the answer inmessage.content; - thinking counts against
max_tokens, so budget generously (2000+). Ifcontentcomes back empty, the model ran out of budget while thinking; raisemax_tokensand retry.
Using it from coding tools
Any tool that accepts a custom OpenAI base URL works:
| Setting | Value |
|---|---|
| Base URL | https://api.green-compute.com/v1 |
| API key | your Green Compute key |
| Model | an id from /v1/models |
Billing
Per token, input and output priced separately, with a small minimum charge per request. See GET /platform/pricing for the live figures.
Green Compute