Distributed Inference on GCP
This repository contains a reproducible deployment for the Alchemyst DevOps internship assignment. It runs the provided quickstart iii project as a small distributed inference mesh on Google Cloud: a public API gateway accepts JSON requests, a TypeScript worker orchestrates the request, and a Python worker runs the GGUF model inference.
[!IMPORTANT] The deployment invariant is that only the API gateway has a public endpoint. The custom workers have no external IP addresses and connect to the iii engine over private subnet networking.
Architecture
Request flow:
nginxreceives the public HTTP request on the gateway.iii-httpinvokeshttp::run_inference_over_http.- The TypeScript caller invokes
inference::get_response. - The caller invokes Python
inference::run_inferencethrough iii RPC. - The Python worker returns
{ "text": "..." }and the gateway sends it as JSON.
API
Endpoint:
POST /v1/chat/completions
Content-Type: application/json
Request:
{
"messages": [
{
"role": "user",
"content": "Say hello in one sentence."
}
]
}
Response:
{
"text": "Hello! I am ready to help."
}
Curl:
curl -X POST "http://$(terraform -chdir=infra/terraform output -raw api_ip)/v1/chat/completions" \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Say hello in one sentence."}]}'
Repository Layout
quickstart/
config.yaml # local iii config with relative worker paths
workers/caller-worker/ # TypeScript HTTP/RPC worker
workers/inference-worker/ # Python model worker
infra/terraform/
main.tf # VPC, subnet, NAT, firewall rules, VMs
variables.tf
outputs.tf
deploy/
gateway/ # iii engine config and nginx reverse proxy
scripts/ # Compute Engine startup scripts
systemd/ # one service per runtime process
Deployment Guides
Use a throwaway GCP project for deployment testing, then destroy it immediately after collecting smoke-test evidence.
The deployment guide covers Google Cloud CLI installation, Terraform
installation, gcloud login, throwaway project creation, API enablement,
Terraform deploy, service checks, API smoke tests, and network checks.
The teardown guide covers terraform destroy, verification that named
resources are gone, manual deletion commands if Terraform cleanup fails,
billing unlink, and project deletion.
Network Checks
Expected properties after terraform apply:
caller-workerandinference-workerhave no external IP addresses.- Public HTTP to the gateway on port
80succeeds. - Public access to gateway port
49134fails. - Worker VMs can reach
10.10.0.10:49134over the private subnet. - SSH access is available only through IAP source range
35.235.240.0/20.
Useful checks:
gcloud compute instances list
gcloud compute ssh alchemyst-devops-caller-worker \
--tunnel-through-iap \
--zone us-central1-a \
--command "timeout 5 bash -c '</dev/tcp/10.10.0.10/49134' && echo reachable"
Local Verification
The lightweight checks do not start the model:
cd quickstart/workers/caller-worker
npm install
npm run build
cd ../inference-worker
python3 -m py_compile inference_worker.py
End-to-end local inference requires iii, the Python dependencies, and the model download:
cd quickstart
iii --config config.yaml
Then run both workers with III_URL=ws://localhost:49134 and send the same curl request to http://localhost:3111/v1/chat/completions.
Production Hardening
Before production, I would add:
- TLS termination with a managed certificate or Caddy.
- Authentication and rate limiting on the public API.
- Request body limits and generation timeouts tuned from real latency data.
- Structured logs, metrics, and traces exported outside the VM.
- Prebuilt images instead of dependency installation in startup scripts.
- Least-privilege service accounts instead of broad default VM permissions.
- Secret Manager for secrets and non-public runtime configuration.
- Dependency lock discipline for Python packages.
If The Model Were 100x Larger
The main change would be to treat inference as a dedicated serving tier instead of a Python worker running raw transformers.generate on CPU. I would move inference to GPU instances or GKE GPU nodes, pre-bake or pre-mount model weights, and use a model server such as vLLM or Text Generation Inference. The API/caller tier would stay lightweight and scale separately.
I would also add queueing, backpressure, and streaming responses so large-model latency does not tie up request threads. The network rule stays the same: model workers remain private, and only the gateway is public.