DeepSeek-R1 has revolutionized open-weights AI reasoning. However, building a commercial SaaS on top of reasoning models requires handling two distinct token streams: internal chain-of-thought (reasoning_content) and final answers. Here is the production architecture using FastAPI and Server-Sent Events (SSE).
1. Dual-Stream Event Payload Architecture
Instead of dumping raw concatenated text into a single response, structured Server-Sent Events allow your Next.js or React frontend to render collapsible thinking accordions in real time while maintaining snappy user feedback.
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
import httpx, json
app = FastAPI()
async def stream_reasoning_generator(prompt: str):
async with httpx.AsyncClient(timeout=60.0) as client:
# Stream from DeepSeek reasoning endpoint
async for chunk in client.stream("POST", "..."):
yield f"data: {json.dumps(chunk)}\n\n"
@app.post("/api/chat/stream")
async def chat_stream(request: ChatRequest):
return StreamingResponse(
stream_reasoning_generator(request.prompt),
media_type="text/event-stream"
)
2. Asynchronous Connection Pooling & VPS Optimization
Reasoning requests often require 15 to 35 seconds of intensive token generation. By employing non-blocking asynchronous generators and HTTPX connection pooling, a single 2-core VPS can sustain hundreds of concurrent client streams without memory exhaustion.