Est.
FeaturesLong read

MCP Hosting Providers That Support Streamable HTTP

Stateless protocol design eliminates session constraints for simpler hosting.

Contributing Editor · · 10 min read
Cover illustration for “MCP Hosting Providers That Support Streamable HTTP”
Features · September 30, 2026 · 10 min read · 2,275 words

The Model Context Protocol's transport layer got rebuilt twice in a little over a year, and each rebuild removed a constraint that used to dictate how MCP servers were hosted. The spec dated 2025-03-26 introduced Streamable HTTP, replacing the old two-endpoint HTTP+SSE model with a single endpoint where the server decides request by request to send back plain JSON or upgrade to an event stream. The 2026-07-28 update went further still. It dropped the protocol-level initialize/initialized handshake, retired the Mcp-Session-Id header, and made the protocol core stateless, so a request can land on any server instance without needing a shared session store or sticky routing to find its way back to the right one.

That single change collapses a whole category of infrastructure that used to be mandatory. A remote MCP server that once required sticky sessions, a shared session store, and deep packet inspection at the gateway can now sit behind a plain round-robin load balancer and route purely on an Mcp-Method header. STDIO, which is how most developers first run an MCP server locally, is a subprocess model tied to a single process on a single machine, and concurrent load causes it to fail catastrophically because it cannot be exposed remotely or scaled out. Streamable HTTP was built specifically to escape that ceiling.

So what does a team actually gain from all this? The architectural question changes from "how do I keep a transport session alive across requests" to "how do I run compute efficiently, and where does state live if my application genuinely needs some," and that's a materially easier problem, but only if the hosting platform's own runtime model doesn't reintroduce the constraints the protocol just removed. Statelessness unlocks serverless and edge deployment in theory. Whether a given platform actually delivers on that in practice depends on how it handles routing, streaming, cold starts, and auth, which is the real evaluation the next section works through.

The four infrastructure properties that separate capable hosts from merely adequate ones

Raw compute specs, vCPU counts, memory tiers, regional footprint, don't tell you whether a platform fits Streamable HTTP MCP hosting. Four properties do: stateless request routing, streaming response handling, cold-start behavior, and authentication support. Each one maps to a way teams get burned by picking the wrong platform, not a theoretical checklist item.

Start with routing. The current MCP spec carries no transport session to persist, so any platform that still forces sticky routing or a shared session store is solving a problem that no longer exists, and paying the cost and complexity of that solution for nothing. Platforms built to route on HTTP headers and scale horizontally without session affinity are the ones actually aligned with where the spec landed. That's a real filter: it separates hosts built around the assumption of persistent connections from hosts built around plain, stateless HTTP semantics.

Streaming is the second filter, and it's where teams get quietly burned. The server still has to be able to upgrade a response to an SSE stream whenever a request calls for it, and a platform with hard response-size limits or a habit of killing long-lived connections will silently break that streaming behavior. The failure appears only under production traffic, not in a load test.

Cold starts matter for a related but distinct reason. Stateless MCP servers scale to zero between requests by design, and cold-start latency directly shapes the experience of every tool call that happens to land on a cold instance. That behavior is a function of language runtime, dependency footprint, container image size, and how much initialization work the server does on boot.

Then there's authentication, the one criterion with a genuine security cost attached to getting it wrong. Running an MCP server on the open internet without auth is a real exposure, and OAuth 2.1 is where the ecosystem has converged as the standard. A platform that ships built-in OAuth integration or first-party auth tooling saves a team from building and maintaining a security-critical layer by hand, which is not a small thing to hand off. These four properties are what the following sections actually test against each provider category.

Serverless and edge platforms suited to stateless Streamable HTTP

Given a genuinely stateless protocol, serverless and edge runtimes become the most natural fit architecturally: no persistent connection to maintain, horizontal scaling that costs nothing extra, and idle compute that costs nothing at all. Two platforms in this tier illustrate the range of what "serverless-native" can mean in practice, and they aren't really competing with each other so much as serving different starting points.

One edge-compute platform in this category ships first-party MCP tooling that addresses the two hardest parts of remote deployment directly: session state, handled through Durable Objects for the cases where an application genuinely needs to remember something across requests, and OAuth, with native Streamable HTTP support built into the platform's workers model. It added Python and JavaScript RPC interoperability in August 2026, widening the languages a team can build in without leaving the platform's own execution model. The free tier is unusually complete for a free tier: unlimited bandwidth, DDoS protection, DNS, and a basic managed WAF ruleset with five custom rules, with full OWASP coverage and the broader managed ruleset set gated behind Pro. Security here isn't something unlocked only once a team starts paying. The default tier's timeout policy will close an SSE stream before a genuinely long-lived streaming session finishes, so anything that needs to hold a connection open for a while requires Durable Objects or a stateful backend layered on top of the stateless default. And the platform leans CPU, not GPU. Tools that need a wide selection of dedicated GPU hardware are better served by the platform covered in the next section.

The other platform in this tier takes a different entry point: a framework-first web platform built around Next.js, where MCP routes sit inside the same deployment as an existing web application rather than running as a separate standalone server. Its Fluid Compute model improves how resources get shared and how concurrency is handled across serverless invocations. Its free plan is a genuinely usable always-on option among major serverless platforms, with a meaningful monthly allotment of function invocations at no cost. Function timeout limits vary by plan and have been revised upward from what a widely used developer guide originally reported, so teams should check current plan limits directly rather than assume a fixed ceiling from older documentation. The fit here is specific: teams that already have a Next.js application and want MCP routes living inside that same deployment unit, not teams looking to stand up an independent MCP service.

AI-native infrastructure for compute-intensive and GPU-dependent MCP tools

Not every MCP server is a thin wrapper around an API call or a database query. Some MCP servers run actual AI inference, execute AI-generated code, or handle untrusted workloads, and those need an execution environment built for that reality rather than just a faster container.

One platform built specifically for this tier describes itself as AI-native serverless compute, and it publishes a first-party example for deploying a remote, stateless Streamable HTTP MCP server using FastMCP. The development model splits cleanly: developers write the application and declare its dependencies in Python, expose the MCP server as a web endpoint over Streamable HTTP, and let the platform handle runtime provisioning, while retaining fine-grained control over CPU, memory, GPU selection, container images, secrets, region, and concurrency. Billing runs per second, so cost tracks the compute a request actually consumes instead of requiring a server to sit running around the clock.

For agents executing untrusted code, the platform offers Sandboxes, a gVisor-isolated runtime kept separate from the MCP process itself, with the next-generation V2 Sandbox backend recommended for high-concurrency workloads. Cold starts get engineered down through an optimized filesystem and memory snapshotting, shortening the initialization-heavy startup path that would otherwise dominate latency for GPU workloads. GPU access spans a wide range, from T4 up through B300, with CPU and GPU work running on the same platform rather than requiring a separate provider for each. It has completed a SOC 2 Type II audit and supports HIPAA-compliant workloads on Enterprise plans via a signed BAA, relevant for teams in regulated industries.

None of this matters if the MCP server in question is a lightweight lookup tool. It matters for servers running inference, doing browser automation, executing generated code in isolation, or needing GPU selection that a general-purpose serverless platform simply doesn't offer. The fit has to match the workload.

Managed container and cloud-native paths for teams with existing AWS or Azure infrastructure

Plenty of teams aren't choosing a hosting platform from scratch. They're already running inside AWS or Azure, under an existing security perimeter and compliance posture, and the lowest-friction path for them runs through the managed services they already operate rather than adding a new vendor to the stack.

On AWS, a purpose-built option now exists: AgentCore Runtime, launched specifically for MCP server deployment, handles container orchestration and scaling automatically rather than requiring a team to wire that up manually. For teams with simpler, genuinely stateless servers, Lambda behind API Gateway also works, using the documented AWS Lambda Web Adapter pattern to bridge a standard server into Lambda's invocation model. This path fits teams whose MCP server needs to sit inside an existing AWS security perimeter, IAM policy boundary, or compliance regime, where standing up a separate platform would mean duplicating controls that already exist.

Azure's answer sits in Azure Functions, which ships a dedicated MCP hosting tutorial in its official documentation and supports Streamable HTTP directly. Its built-in OAuth integration cuts down the authentication boilerplate, one of the four evaluation criteria from earlier in this piece, directly at the platform level. That fits especially well for enterprises already running Azure Active Directory as their identity provider, since MCP server auth can then unify with IdP configuration that already exists rather than standing up a parallel auth system.

None of this comes free of trade-offs. Cloud-native paths carry more configuration surface and slower iteration cycles than a PaaS deployment would, but for teams already operating inside these environments, the operational familiarity and compliance inheritance outweigh the extra setup cost.

PaaS options for conventional Python and Node.js servers that need minimal infrastructure overhead

Somewhere between full cloud-native and pure serverless is a large group of teams whose local Python or Node.js MCP server already runs and works, and all it actually needs is a Streamable HTTP adapter, an auth layer, and a public URL. PaaS platforms are built for exactly that migration, and they trade a bit of per-unit cost for a meaningful drop in operational overhead.

One option in this tier takes the lowest-friction route for a conventional repository: push to GitHub, and the platform builds and deploys it. That's a real advantage when the MCP server isn't standalone, when it needs to call several internal services, or needs a Redis instance, a database, a queue, or a background worker running in the same project. SSE works without any special configuration because Railway runs containers as long-lived processes rather than through serverless invocation. Cost stays modest for the common case: a small MCP server running on 512 MB of RAM and minimal CPU costs a low monthly price including bandwidth, based on figures from a widely used developer guide and separate research.

A second option offers fixed-instance predictability instead of usage-based pricing, with Web Service plans at 512 MB and 2 GB RAM tiers priced by memory allocation. It handles SSE and WebSockets correctly, and publishes official MCP server templates for both Python (FastMCP) and TypeScript (the official TypeScript SDK), each including a render.yaml Blueprint, Streamable HTTP transport, a health check endpoint, and an auto-generated bearer token. Its free tier spins down after a period of inactivity, which makes it a fine place to prototype but not somewhere to run anything meant to stay reachable. The fit for both platforms in this tier is the same in spirit: teams that want a managed web service model with clear instance boundaries and no interest in hand-managing systemd or TLS certificates themselves.

Self-managed VPS as the baseline for teams with volume, control requirements, or tight budgets

Most production MCP servers running at modest scale still sit on a small VPS behind a reverse proxy, and that's not because it's the architecturally cleanest answer. It persists because it's simple, predictable, and cheap for a server with stable, foreseeable traffic. None of the managed convenience from the PaaS or serverless tiers comes with this option, but none of their constraints do either.

The self-managed model treats stateless and stateful SSE servers the same way. There's no platform timeout policy to work around, no invocation limit to budget against, and full control over the runtime and the network configuration. For a team with steady, well-understood traffic and someone on staff who can own the operational side, that trade often works out fine.

The standard setup provisions a VPS, installs Node or Python, clones the MCP server repository, wraps the running process in systemd or PM2 so it restarts on failure, and puts Caddy or nginx in front for routing, none of it exotic. It's also entirely manual, and every piece of it, the certificate renewal, the process supervision, the security patching, the failover, is now the team's responsibility rather than a platform's. For teams with the volume to make that worthwhile or the budget constraints that rule out a managed tier, it's a legitimate baseline. For teams without spare operational capacity, the PaaS and serverless tiers covered earlier exist precisely to take that weight off the table.

Sources

  1. Best Platforms for Streamable HTTP MCP Servers in 2026 | Modal Blog
  2. Building and hosting MCP servers: a complete guide
  3. MCP Server Hosting: Where to Actually Deploy in 2026 - DEV Community
  4. The 2026-07-28 Specification | Model Context Protocol Blog
  5. Streamable HTTP - Model Context Protocol
  6. Cloudflare MCP servers support the new MCP 2026-07-28 Specification · Changelog
  7. HTTP Deployment - FastMCP

More in Features