Running OpenAI's gpt-oss-20b with 128k Context on a Single L4 GPU
Alexey Nizhegolenko DevOps Engineer, AgentOps Engineer, AI Infrastructure Engineer This is the second article in my series on self-hosting LLMs on GKE. In the first article I covered deploying Gemma4 26B with a 28,000 token context window. This time I'll show you something more impressive: openai/gpt-oss-20b running with a 128,000 token context on the same single L4 GPU. The setup has been running in production since November 2025, for about 6 months, with no major incidents. That's the kind...
Original Source
Read the full article at Dev →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.