A minimal but real demonstration of deploying a production-style AI/API service inside a Virtual Private Cloud, with the app fully network-isolated and reachable only through a load balancer. Built to demonstrate enterprise-style network security patterns for AI system deployment.
- VPC design with public and private subnets
- An application server with no direct internet exposure — it only accepts traffic from an internal load balancer
- Security group / firewall rules that enforce least-privilege network access
- A working FastAPI service (RAG-style
/askendpoint) deployed inside that private network - Infrastructure defined as code (Terraform), not clicked together by hand
flowchart TB
Internet((Internet))
subgraph VPC["VPC: ai-service-vpc"]
subgraph PublicSubnet["Public Subnet (10.0.1.0/24)"]
ALB["Application Load Balancer"]
end
subgraph PrivateSubnet["Private Subnet (10.0.2.0/24)"]
App["FastAPI App Server\n(no public IP)"]
VDB["Vector DB / Cache\n(optional, no public access)"]
end
NAT["Cloud NAT\n(outbound only)"]
Router["Cloud Router"]
end
LLM[("External LLM API")]
Internet -->|HTTPS| ALB
ALB -->|"tcp:8000, allow-lb-to-app rule"| App
App --> VDB
App -.->|outbound only| Router
Router --> NAT
NAT -.->|outbound only| LLM
style App fill:#1f2937,stroke:#22c55e,color:#fff
style ALB fill:#1f2937,stroke:#3b82f6,color:#fff
style NAT fill:#1f2937,stroke:#f59e0b,color:#fff
Key principle: the app server has no public IP and no route directly to/from the internet. The only way in is through the load balancer, and the only way the app reaches the outside world (e.g., calling an LLM API) is through a NAT Gateway — outbound only, no inbound.
Enterprise AI deployments increasingly handle sensitive data — user queries, retrieved documents, proprietary context fed into prompts — and that data has to move through infrastructure that doesn't casually expose it. Putting the compute layer in a private subnet with no public IP means the attack surface for the app itself is reduced to exactly one path: through a load balancer that can be monitored, rate-limited, and put behind WAF rules. Even if the load balancer is compromised or misconfigured, there's no direct route to the machine running the model logic. Outbound traffic (calls to an LLM provider, telemetry, etc.) is deliberately restricted to NAT-only egress, so the service can call out to what it needs without ever being reachable from outside. This is the same shape most compliance frameworks (SOC 2, HIPAA-adjacent workloads) expect to see for anything touching user data — it's not exotic, it's the baseline enterprises look for before trusting a vendor's AI service with their traffic.
app/ FastAPI application (LLM adapter, vector store, endpoints)
infra/ Terraform: VPC, subnets, firewall, NAT, load balancer, compute instance
docs/ Firewall rule reference, isolation verification steps, screenshots
Dockerfile Container build for the app
docker-compose.yml Local dev: run the app + vector store together
- Cloud provider: GCP (Free Tier)
- App: FastAPI (Python), provider-agnostic LLM adapter (Anthropic / OpenAI / local echo mode)
- Vector store: Chroma (embedded, no separate managed service needed)
- Networking: VPC, subnets, firewall rules, Cloud Load Balancing, Cloud NAT — via Terraform
- Compute: A single
e2-smallCompute Engine VM, no public IP
Test the actual application logic before touching any cloud console.
cp .env.example .env
# Leave LLM_PROVIDER=none to run fully offline (echo mode), or set it to
# "anthropic"/"openai" and paste your own key into LLM_API_KEY in .env.
docker compose up --buildThen:
curl http://localhost:8000/health
curl -X POST http://localhost:8000/ask \
-H "Content-Type: application/json" \
-d '{"question": "What is this service?"}'cd infra
cp terraform.tfvars.example terraform.tfvars
# Fill in your project_id, region, and admin_ip (your own public IP for SSH access)
terraform init
terraform plan
terraform applyThis stands up the VPC, both subnets, firewall rules, Cloud Router + NAT, the load balancer chain, an Artifact Registry repository, a Secret Manager secret container, and the compute instance definition. See infra/terraform.tfvars.example for exactly which values you need to supply — no real credentials or project IDs are committed to this repo.
Where your API key goes on the real VM: Terraform creates an empty Secret Manager secret and grants the VM read access to it, but never sets the actual value — you add that yourself with one gcloud secrets versions add command after applying. Full steps in docs/secret-manager.md.
The compute instance pulls its container image from Artifact Registry at boot (var.app_image_tag). Build and push it before or after the first terraform apply — full commands in docs/artifact-registry.md, short version:
gcloud builds submit --tag us-central1-docker.pkg.dev/<project-id>/ai-service-repo/ai-service:latest .Then set app_image_tag in terraform.tfvars to that same value and re-apply.
Full steps in docs/verification.md. The short version:
# From your local machine — this should FAIL (no public IP, no route)
curl http://<private-instance-internal-ip>:8000/health
# Through the load balancer — this should SUCCEED
curl http://<load-balancer-ip>/healthThat contrast is the actual proof point for "VPC deployment" experience — see docs/screenshots/ for what to capture once deployed.
See docs/firewall-rules.md for the full rule-by-rule breakdown of what's allowed and why.
Everything here fits inside GCP's Free Tier (e2-small VM, low NAT usage, low LB traffic) if torn down after documenting it. An ALB and a running VM left on 24/7 will eventually cost a few dollars a month.
cd infra
terraform destroyConfirm in the GCP Console that the VM, load balancer, NAT, and subnets are actually gone afterward — terraform destroy should catch everything defined here, but it's worth checking for stray resources (e.g. reserved IPs) that can otherwise keep billing.
MIT — see LICENSE.