Two years ago we had a client running its embedding pipeline on AWS and its fine-tuned inference endpoints on Azure OpenAI, because the fine-tuning quota they needed only existed on Azure at the time. Every request crossed the public internet twice — once to fetch context, once to get the prediction back — through a VPN tunnel one of our engineers built and nobody wanted to touch again. It worked. It was also the flakiest part of the entire stack, and every postmortem for six months had a line that said “packet loss on the cross-cloud hop.”
That’s the exact shape of problem AWS Interconnect – multicloud with Azure is built to remove, and it went into public preview quietly in August. I spent an afternoon reading through the architecture instead of the press release, because the press release undersells what’s actually interesting here.
What it is, specifically
A private, provisioned connection between an AWS region and an Azure region, set up through the console, CLI, or Azure portal in a few clicks instead of the weeks-long process of arranging your own cross-connect at a colocation facility. Every link runs IEEE 802.1AE MACsec encryption between the AWS and Azure edge routers — hardware-level encryption on the physical hop, not something layered on at the application level. Each connection provisions four independent logical paths across physically separate interconnect facilities and routers, so losing one site or doing router maintenance on one path doesn’t take the connection down.
Available today in four regions: US East (N. Virginia), US West (N. California), Asia Pacific (Sydney), and Europe (Frankfurt). AWS is quoting a 99.99% availability target — explicitly a design goal during preview, not a contractual SLA yet.
# Rough shape of provisioning via AWS CLI (illustrative — check current API for exact params)
aws directconnect create-interconnect-multicloud \
--partner-cloud azure \
--azure-region eastus \
--bandwidth 1Gbps \
--region us-east-1
The number that actually matters: 1 Gbps
Preview partners are capped at 1 gigabit. Compare that to the free 500 Mbps local interconnect AWS already offers to generally-available partners — a tier Azure hasn’t reached yet. If you’re picturing this as a pipe for shipping training data or moving multi-terabyte checkpoints between clouds, it isn’t that, not yet, not at this bandwidth.
What it is good for: request-response traffic. Inference calls. RAG lookups that fetch a vector match from one cloud and generate the answer on the other. Metadata and control-plane traffic between services that happen to live on different clouds because of a vendor quota, an acquisition, or a compliance requirement that pinned one workload to a specific provider. That’s a narrower use case than “multicloud networking” sounds like, and it’s also the exact use case that was causing us pain.
Why this is an AI infrastructure story, not just a networking one
Multi-model AI stacks are multi-vendor by default now, whether teams planned it that way or not. You fine-tune where the quota and the base model you want are both available. You run vector search wherever your existing data warehouse already lives. You call three different providers’ APIs from the same orchestration layer because no single vendor covers embeddings, generation, and evaluation equally well. That fragmentation used to mean every cross-cloud call went over the public internet, through egress fees on both sides, with latency and packet loss you couldn’t do anything about except retry harder.
A private, redundant, encrypted-by-default link between AWS and Azure — the two clouds most enterprise AI workloads actually straddle — turns “our RAG pipeline spans two cloud providers” from an architectural liability into a boring implementation detail. That’s a real shift, even capped at 1 Gbps, because for inference-shaped traffic 1 Gbps is not the bottleneck. Application-level retry logic and DNS resolution across the public internet were the bottleneck, and those go away.
What I’d actually do with this today
I would not migrate a production workload onto a service that’s explicitly in preview with a target-not-guaranteed SLA. That’s an easy trap: the demo works, the four-path redundancy story sounds solid, and six weeks later you’ve built a dependency on something AWS could still change the shape of before GA.
What I would do — and what we’re doing with the client from the story above — is set up the interconnect in a staging environment now, in one of the four supported regions, and start routing non-critical inference traffic across it. Measure the actual latency and packet loss against the VPN tunnel it would replace. Build the failure-mode runbook before there’s an incident that forces you to write one under pressure. When it hits general availability with a real SLA and Azure’s own 500 Mbps free tier extends to it, you want to already know whether the answer is “migrate everything” or “this only helps for the RAG lookups, keep the VPN for the rest.”
The most expensive mistake in multicloud isn’t picking the wrong two providers. It’s discovering the shape of your actual cross-cloud traffic pattern for the first time during an incident, instead of during a quiet Tuesday afternoon with a staging environment and nothing on fire.