Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

VPC-Isolated AI Service Deployment

A minimal but real demonstration of deploying a production-style AI/API service inside a Virtual Private Cloud, with the app fully network-isolated and reachable only through a load balancer. Built to demonstrate enterprise-style network security patterns for AI system deployment.

What this project demonstrates

  • VPC design with public and private subnets
  • An application server with no direct internet exposure — it only accepts traffic from an internal load balancer
  • Security group / firewall rules that enforce least-privilege network access
  • A working FastAPI service (RAG-style /ask endpoint) deployed inside that private network
  • Infrastructure defined as code (Terraform), not clicked together by hand

Architecture

flowchart TB
    Internet((Internet))

    subgraph VPC["VPC: ai-service-vpc"]
        subgraph PublicSubnet["Public Subnet (10.0.1.0/24)"]
            ALB["Application Load Balancer"]
        end

        subgraph PrivateSubnet["Private Subnet (10.0.2.0/24)"]
            App["FastAPI App Server\n(no public IP)"]
            VDB["Vector DB / Cache\n(optional, no public access)"]
        end

        NAT["Cloud NAT\n(outbound only)"]
        Router["Cloud Router"]
    end

    LLM[("External LLM API")]

    Internet -->|HTTPS| ALB
    ALB -->|"tcp:8000, allow-lb-to-app rule"| App
    App --> VDB
    App -.->|outbound only| Router
    Router --> NAT
    NAT -.->|outbound only| LLM

    style App fill:#1f2937,stroke:#22c55e,color:#fff
    style ALB fill:#1f2937,stroke:#3b82f6,color:#fff
    style NAT fill:#1f2937,stroke:#f59e0b,color:#fff
Loading

Key principle: the app server has no public IP and no route directly to/from the internet. The only way in is through the load balancer, and the only way the app reaches the outside world (e.g., calling an LLM API) is through a NAT Gateway — outbound only, no inbound.

Why this pattern matters

Enterprise AI deployments increasingly handle sensitive data — user queries, retrieved documents, proprietary context fed into prompts — and that data has to move through infrastructure that doesn't casually expose it. Putting the compute layer in a private subnet with no public IP means the attack surface for the app itself is reduced to exactly one path: through a load balancer that can be monitored, rate-limited, and put behind WAF rules. Even if the load balancer is compromised or misconfigured, there's no direct route to the machine running the model logic. Outbound traffic (calls to an LLM provider, telemetry, etc.) is deliberately restricted to NAT-only egress, so the service can call out to what it needs without ever being reachable from outside. This is the same shape most compliance frameworks (SOC 2, HIPAA-adjacent workloads) expect to see for anything touching user data — it's not exotic, it's the baseline enterprises look for before trusting a vendor's AI service with their traffic.

Repo layout

app/            FastAPI application (LLM adapter, vector store, endpoints)
infra/          Terraform: VPC, subnets, firewall, NAT, load balancer, compute instance
docs/           Firewall rule reference, isolation verification steps, screenshots
Dockerfile      Container build for the app
docker-compose.yml   Local dev: run the app + vector store together

Stack

  • Cloud provider: GCP (Free Tier)
  • App: FastAPI (Python), provider-agnostic LLM adapter (Anthropic / OpenAI / local echo mode)
  • Vector store: Chroma (embedded, no separate managed service needed)
  • Networking: VPC, subnets, firewall rules, Cloud Load Balancing, Cloud NAT — via Terraform
  • Compute: A single e2-small Compute Engine VM, no public IP

Running it locally first

Test the actual application logic before touching any cloud console.

cp .env.example .env
# Leave LLM_PROVIDER=none to run fully offline (echo mode), or set it to
# "anthropic"/"openai" and paste your own key into LLM_API_KEY in .env.

docker compose up --build

Then:

curl http://localhost:8000/health
curl -X POST http://localhost:8000/ask \
  -H "Content-Type: application/json" \
  -d '{"question": "What is this service?"}'

Deploying the infrastructure

cd infra
cp terraform.tfvars.example terraform.tfvars
# Fill in your project_id, region, and admin_ip (your own public IP for SSH access)

terraform init
terraform plan
terraform apply

This stands up the VPC, both subnets, firewall rules, Cloud Router + NAT, the load balancer chain, an Artifact Registry repository, a Secret Manager secret container, and the compute instance definition. See infra/terraform.tfvars.example for exactly which values you need to supply — no real credentials or project IDs are committed to this repo.

Where your API key goes on the real VM: Terraform creates an empty Secret Manager secret and grants the VM read access to it, but never sets the actual value — you add that yourself with one gcloud secrets versions add command after applying. Full steps in docs/secret-manager.md.

Building and pushing the app image

The compute instance pulls its container image from Artifact Registry at boot (var.app_image_tag). Build and push it before or after the first terraform apply — full commands in docs/artifact-registry.md, short version:

gcloud builds submit --tag us-central1-docker.pkg.dev/<project-id>/ai-service-repo/ai-service:latest .

Then set app_image_tag in terraform.tfvars to that same value and re-apply.

Verifying isolation

Full steps in docs/verification.md. The short version:

# From your local machine — this should FAIL (no public IP, no route)
curl http://<private-instance-internal-ip>:8000/health

# Through the load balancer — this should SUCCEED
curl http://<load-balancer-ip>/health

That contrast is the actual proof point for "VPC deployment" experience — see docs/screenshots/ for what to capture once deployed.

Firewall rules

See docs/firewall-rules.md for the full rule-by-rule breakdown of what's allowed and why.

Cost note

Everything here fits inside GCP's Free Tier (e2-small VM, low NAT usage, low LB traffic) if torn down after documenting it. An ALB and a running VM left on 24/7 will eventually cost a few dollars a month.

Teardown

cd infra
terraform destroy

Confirm in the GCP Console that the VM, load balancer, NAT, and subnets are actually gone afterward — terraform destroy should catch everything defined here, but it's worth checking for stray resources (e.g. reserved IPs) that can otherwise keep billing.

License

MIT — see LICENSE.

About

No description or website provided.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages