OP 09 September, 2026 - 03:25 PM
(This post was last modified: 19 September, 2026 - 06:46 PM by Talmud. Edited 2 times in total.)
UPDATED / 9/19/2026
Currently, there is an abundance of "AI-as-a-Service" hype. I've been looking into the backends of hermres, fx.sh, and, most amusingly, OpenCode.ai, and let me tell you, these platforms are far from production-ready. They are leaking free inference like a sieve, and it is trivially simple to steal "AI compute".
This thread discusses OpenCode.ai and its public API endpoint. Everything here is based on what I observed in the wild. There are no opinions, just raw mechanics.
The endpoint URL is https://opencode.ai/zen/v1/responses. It responds to POST requests with a predefined set of headers and a JSON body. No API key means no sign-up. If you use the correct headers, you will get Muse model outputs.
The most powerful free models available in OpenCode are muse-spark-1.2-contributor-free and muse-spark-1.3-contributor-free. Others exist, but these two are the best for general use in my opinion.
The required headers are:
The session and request IDs are generated locally. Here is how they are built: take the current timestamp in milliseconds, multiply by 4096, add a per‑millisecond counter, then mask to 48 bits. For the session ID, you invert that masked value; for the request ID, you keep it as is. Then format that as 12 hex digits, append 14 random characters from 0‑9A‑Za‑z, and prefix with "ses_" or "msg_". The server does not check if these are unique or tied to anything; it only checks the format.
The JSON payload is now built from a template file called share_template.json. That file contains the base request body. The code loads it, then replaces the model, input, and other fields. The input array must contain at least two elements: a system message (index 0) and a user message (index 1). The code overwrites index 1 with your prompt. If you provide a system prompt, it is prepended to the user text with two newlines. The payload also includes a prompt_cache_key set to the session ID, and optionally max_output_tokens.
When you get a response, it may be JSON or SSE (server-sent events). The new code parses both. For SSE, it looks for lines starting with "data: ", then checks the type. It accumulates text from "response.output_text.delta" events and marks completion on "response.completed". For JSON, it extracts text from the "output" array, looking for items with "type": "output_text". If that fails, it falls back to the top‑level "output_text" field.
Here is the updated code lifted directly from the wild:
What changed from the previous version:
The User-Agent string updated to opencode/1.18.30 ai-sdk/provider-utils/4.0.40 runtime/bun/1.3.14.
A new header Accept: */* and x-opencode-project: global were added.
The payload is now built from a template file share_template.json instead of being hardcoded. The code loads it once and caches it > which you can download with the script HERE
The payload structure changed: instead of a simple "input": "prompt", it now uses an array of message objects. The code places your prompt into the second element (index 1) as a user message, optionally prepending a system prompt.
A prompt_cache_key is set to the session ID.
The ID generation is now thread-safe with a lock.
The HTTP request now explicitly uses an opener and reads the raw response as text, because the response may be SSE.
A new _parse_sse function handles server-sent events, accumulating text from response.output_text.delta events.
The _extract_text function now filters by message role (only assistant or no role) and only accepts output_text type.
The muse function signature expanded to accept project_id, parent_session_id, extra_headers, and api_key.
The main block now outputs structured JSON logs (step_start, text, step_finish, error) for integration with other tools.
Currently, there is an abundance of "AI-as-a-Service" hype. I've been looking into the backends of hermres, fx.sh, and, most amusingly, OpenCode.ai, and let me tell you, these platforms are far from production-ready. They are leaking free inference like a sieve, and it is trivially simple to steal "AI compute".
This thread discusses OpenCode.ai and its public API endpoint. Everything here is based on what I observed in the wild. There are no opinions, just raw mechanics.
The endpoint URL is https://opencode.ai/zen/v1/responses. It responds to POST requests with a predefined set of headers and a JSON body. No API key means no sign-up. If you use the correct headers, you will get Muse model outputs.
The most powerful free models available in OpenCode are muse-spark-1.2-contributor-free and muse-spark-1.3-contributor-free. Others exist, but these two are the best for general use in my opinion.
The required headers are:
Code:
Authorization: Bearer public
Content-Type: application/json
Accept: */*
User-Agent: opencode/1.18.30 ai-sdk/provider-utils/4.0.40 runtime/bun/1.3.14
x-opencode-session: ses_<12‑char hex><14 random alnum chars>
x-opencode-request: msg_<same format>
x-opencode-client: cli
x-opencode-project: globalThe session and request IDs are generated locally. Here is how they are built: take the current timestamp in milliseconds, multiply by 4096, add a per‑millisecond counter, then mask to 48 bits. For the session ID, you invert that masked value; for the request ID, you keep it as is. Then format that as 12 hex digits, append 14 random characters from 0‑9A‑Za‑z, and prefix with "ses_" or "msg_". The server does not check if these are unique or tied to anything; it only checks the format.
The JSON payload is now built from a template file called share_template.json. That file contains the base request body. The code loads it, then replaces the model, input, and other fields. The input array must contain at least two elements: a system message (index 0) and a user message (index 1). The code overwrites index 1 with your prompt. If you provide a system prompt, it is prepended to the user text with two newlines. The payload also includes a prompt_cache_key set to the session ID, and optionally max_output_tokens.
When you get a response, it may be JSON or SSE (server-sent events). The new code parses both. For SSE, it looks for lines starting with "data: ", then checks the type. It accumulates text from "response.output_text.delta" events and marks completion on "response.completed". For JSON, it extracts text from the "output" array, looking for items with "type": "output_text". If that fails, it falls back to the top‑level "output_text" field.
Here is the updated code lifted directly from the wild:
Code:
import copy
import json
import os
import secrets
import threading
import time
import urllib.error
import urllib.request
BASE = "https://opencode.ai/zen/v1"
USER_AGENT = "opencode/1.18.30 ai-sdk/provider-utils/4.0.40 runtime/bun/1.3.14"
DEFAULT_MODEL = "muse-spark-1.3-contributor-free"
MUSE_FREE = {
"muse-spark-1.2-contributor-free",
"muse-spark-1.3-contributor-free",
}
_TEMPLATE_FILE = os.path.join(os.path.dirname(os.path.abspath(file)), "share_template.json")
_template_cache = None
_lock = threading.Lock()
_counter_ms = 0
_counter_n = 0
_ALPHABET = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"
def _make_id(invert, timestamp_ms=None):
global _counter_ms, _counter_n
with _lock:
if timestamp_ms is None:
timestamp_ms = int(time.time() * 1000)
if timestamp_ms != _counter_ms:
_counter_ms = timestamp_ms
_counter_n = 0
_counter_n += 1
value = timestamp_ms * 0x1000 + _counter_n
mask = (1 << 48) - 1
value = (mask ^ (value & mask)) if invert else (value & mask)
suffix = "".join(_ALPHABET[b % 62] for b in secrets.token_bytes(14))
return f"{value:012x}{suffix}"
SESSION_ID = "ses" + _make_id(True)
def _template():
global _template_cache
if _template_cache is None:
with open(_TEMPLATE_FILE, encoding="utf-8") as f:
_template_cache = json.load(f)
return copy.deepcopy(_template_cache)
def _headers(project_id=None, parent_session_id=None, extra=None, api_key=None):
h = {
"Authorization": f"Bearer {api_key or 'public'}",
"Content-Type": "application/json",
"Accept": "*/*",
"User-Agent": USER_AGENT,
"x-opencode-session": SESSION_ID,
"x-opencode-request": "msg" + _make_id(False),
"x-opencode-client": "cli",
"x-opencode-project": project_id or "global",
}
if parent_session_id:
h["x-parent-session-id"] = parent_session_id
if extra:
h.update(extra)
return h
def _parse_sse(raw):
text = ""
completed = False
for line in raw.splitlines():
if not line.startswith("data: "):
continue
try:
d = json.loads(line[6:])
except Exception:
continue
if d.get("type") == "response.output_text.delta":
text += d.get("delta", "")
elif d.get("type") == "response.completed":
completed = True
return text, completed
def _extract_text(data):
texts = []
for output in data.get("output", []) or []:
if output.get("type") != "message":
continue
if "role" in output and output.get("role") not in ("assistant", None):
continue
for item in output.get("content", []) or []:
if not isinstance(item, dict):
continue
if item.get("type") == "output_text" and isinstance(item.get("text"), str):
texts.append(item["text"])
if texts:
return "".join(texts)
return data.get("output_text", "") or ""
def muse(prompt, system=None, timeout=90, max_tokens=None, model=None, project_id=None, parent_session_id=None, extra_headers=None, api_key=None):
model = model or DEFAULT_MODEL
if model not in MUSE_FREE:
raise ValueError(f"Use {', '.join(sorted(MUSE_FREE))}")
payload = _template()
payload["model"] = model
user_text = (system + "\n\n" + prompt) if system else prompt
payload["input"][1] = {"role": "user", "content": [{"type": "input_text", "text": user_text}]}
payload["prompt_cache_key"] = _SESSION_ID
if max_tokens:
payload["max_output_tokens"] = max_tokens
request = urllib.request.Request(
f"{BASE}/responses",
data=json.dumps(payload).encode(),
headers=_headers(project_id=project_id, parent_session_id=parent_session_id, extra=extra_headers, api_key=api_key),
)
opener = urllib.request.build_opener()
try:
with opener.open(request, timeout=timeout) as response:
raw = response.read().decode("utf-8", "replace")
except urllib.error.HTTPError as exc:
try:
detail = exc.read().decode()[:1500]
except Exception:
detail = ""
raise RuntimeError(f"HTTP {exc.code}: {detail or exc.reason}") from exc
text, completed = _parse_sse(raw)
if not text:
try:
data = json.loads(raw) if raw.strip().startswith("{") else {}
except Exception:
data = {}
text = _extract_text(data)
if not text:
raise RuntimeError("empty text (completed=" + str(completed) + ")")
return text
if name == "main":
import sys
prompt = " ".join(sys.argv[1:]) if len(sys.argv) > 1 else "Say hi in 5 words."
now = lambda: int(time.time() * 1000)
session_id = SESSION_ID
message_id = "msg" + make_id(False)
print(json.dumps({"type": "step_start", "timestamp": now(), "sessionID": session_id, "part": {"id": "prt" + make_id(False), "messageID": message_id, "sessionID": session_id, "type": "step-start"}}))
try:
text = muse(prompt)
t0 = now()
print(json.dumps({"type": "text", "timestamp": t0, "sessionID": session_id, "part": {"id": "prt" + make_id(False), "messageID": message_id, "sessionID": session_id, "type": "text", "text": text, "time": {"start": t0, "end": t0}}}))
print(json.dumps({"type": "step_finish", "timestamp": now(), "sessionID": session_id, "part": {"id": "prt" + _make_id(False), "reason": "stop", "messageID": message_id, "sessionID": session_id, "type": "step-finish"}}))
except Exception as exc:
print(json.dumps({"type": "error", "timestamp": now(), "sessionID": session_id, "error": {"name": type(exc).name, "data": {"message": str(exc)}}}))What changed from the previous version:
The User-Agent string updated to opencode/1.18.30 ai-sdk/provider-utils/4.0.40 runtime/bun/1.3.14.
A new header Accept: */* and x-opencode-project: global were added.
The payload is now built from a template file share_template.json instead of being hardcoded. The code loads it once and caches it > which you can download with the script HERE
The payload structure changed: instead of a simple "input": "prompt", it now uses an array of message objects. The code places your prompt into the second element (index 1) as a user message, optionally prepending a system prompt.
A prompt_cache_key is set to the session ID.
The ID generation is now thread-safe with a lock.
The HTTP request now explicitly uses an opener and reads the raw response as text, because the response may be SSE.
A new _parse_sse function handles server-sent events, accumulating text from response.output_text.delta events.
The _extract_text function now filters by message role (only assistant or no role) and only accepts output_text type.
The muse function signature expanded to accept project_id, parent_session_id, extra_headers, and api_key.
The main block now outputs structured JSON logs (step_start, text, step_finish, error) for integration with other tools.
![[Image: undefined-Imgur.gif]](https://i.ibb.co/BHYKZvCz/undefined-Imgur.gif)
![[Image: aqNFPAg.gif]](https://i.imgur.com/aqNFPAg.gif)
![[Image: AIObanner.gif]](https://i.ibb.co/Jj0XtLWz/AIObanner.gif)