Two years ago a client asked me to put a lightweight recommendation API on Cloudflare Workers. I said no, and the reason was Python. Their data science team owned the model-serving code in Python, our platform team lived in the Workers ecosystem, and bridging the two meant either rewriting the inference logic in JavaScript or standing up a whole separate compute tier just to host a Python process. We went with the separate tier. It worked, it also added a hop, a deploy pipeline, and a bill.

On September 21st Cloudflare shipped Python Workers to general availability, and the detail that actually matters isn’t “Python runs on Workers now” — it’s that it runs with real TCP sockets. Previous attempts at Python-on-the-edge (including Cloudflare’s own earlier beta) were REST-call sandboxes: fine for calling an HTTP API, useless for talking to a real database. This release adds actual socket support, meaning aiomysql and asyncpg work. That’s the difference between a demo and something you’d put in production.

How they actually did it

Python Workers run on Pyodide — Python compiled to WebAssembly. The interesting engineering is in how they got real networking working inside a WASM sandbox that has no native socket syscalls. Cloudflare intercepts Python’s socket-level calls and translates them into the Workers platform’s own connect() API, which already handles raw TCP for other runtimes. So asyncpg thinks it’s opening a normal socket; underneath, that call is getting rerouted through Workers’ connection layer to a real Postgres instance.

On top of that, they shipped workers.asgi and workers.wsgi bridge modules, so FastAPI, Django, and Flask apps deploy close to unmodified. And — this is the part that matters if you’re doing anything AI-adjacent — requests and httpx get transparently routed through Workers’ fetch(), which means packages built on top of those libraries just work. openai’s Python SDK. langchain, plus a new first-party langchain-cloudflare integration package. The mcp package. All running as edge functions, not as a proxy layer in front of a “real” server somewhere.

Cloudflare also got PEP 783 accepted upstream — a new PyEmscripten platform tag for Python packaging. That’s a quieter but more durable signal than the GA announcement itself: running Python compiled to WASM in a serverless sandbox is now something the Python packaging ecosystem officially recognizes as a target platform, not a hack Cloudflare maintains alone.

What this actually buys you

# wrangler.toml
name = "edge-agent-endpoint"
main = "src/entry.py"
compatibility_flags = ["python_workers"]
compatibility_date = "2026-09-21"
# src/entry.py
from fastapi import FastAPI
from openai import AsyncOpenAI
import asyncpg

app = FastAPI()
client = AsyncOpenAI()  # key injected via Workers secret binding

@app.post("/classify")
async def classify(payload: dict):
    conn = await asyncpg.connect(dsn="postgresql://...")
    history = await conn.fetch(
        "select label from classifications where user_id = $1 limit 5", payload["user_id"]
    )
    response = await client.chat.completions.create(
        model="gpt-6-luna",
        messages=[{"role": "user", "content": payload["text"]}],
    )
    await conn.close()
    return {"label": response.choices[0].message.content}

That’s a real endpoint — DB read, LLM call, DB-informed response — running at whatever Cloudflare edge node is closest to the request, with no separate compute tier to provision. For a lightweight classification or routing endpoint sitting in front of a bigger system, that’s genuinely useful: you get global distribution and Workers’ pricing model for something that used to need its own Lambda or container.

Where I’d still say no

I wouldn’t put anything CPU-heavy here. Pyodide is WASM, and WASM Python is still meaningfully slower than CPython for anything numeric — nobody’s running actual model inference inside the Worker itself, and Cloudflare isn’t claiming that. The pattern is edge Worker as a thin, stateless orchestration layer that calls out to inference elsewhere (Workers AI, Bedrock, wherever), not a place to run your own forward pass.

I’d also want to see this in production for a quarter before trusting it with something latency-sensitive and stateful. Socket support inside a WASM sandbox is new enough — GA as of four days ago — that I’d expect edge cases in connection pooling and timeout behavior nobody’s hit yet. The demo-to-production gap on “real socket support in a browser-engine sandbox” is usually wider than the announcement blog post suggests.

But the client I turned down two years ago? That project is exactly the shape this now fits: a thin classification layer, a small Postgres lookup, an LLM call, deployed globally without owning a second compute tier. If they asked again today, the answer changes to “yes, but let’s pilot it on the low-stakes endpoint first.”

Export for reading

Comments