This is a condensed version of Xuhui’s full write-up, which includes the architecture figures, the runtime comparison tables and the alerting cost model. Read the complete article on xuhuizhan5.github.io.
FORGE did not begin with a search for the most powerful cluster. It began with a recurring delivery problem: a scheduled ML pipeline, a memory-sensitive API, an edge workload, and a conventional service all needed reliable infrastructure, promotion, rollback, and observability, but they did not need the same runtime.
How could those workloads share an operating model without forcing each of them into Kubernetes? FORGE answers with contracts around infrastructure intent, immutable artifacts, reviewed deployment state, health, and operator feedback. EKS is one implementation surface inside that system, alongside managed application and edge runtimes.
Three facts kept the design grounded. Demand is concentrated across the eastern and southern United States, so a compact regional foundation is reasonable. The product is growing quickly, so online services and deadline-bound workflows need independent room to scale. And because most surrounding services already run on AWS, reusing its identity, delivery, networking, metrics, and cost controls is simpler than introducing another hosting control plane.
Standardize contracts, not every runtime
What should these workloads actually share? The responsibilities that operators repeat, not necessarily the scheduler that runs them. FORGE defines six shared contracts: infrastructure declaration, artifact production, deployment-state review, reconciliation, health and metrics, and operator notification. A workload can implement those as a scheduled Kubernetes job, a memory-sensitive API, an edge component, or a conventional managed service. That keeps EKS in its proper place: useful where Kubernetes scheduling creates leverage, but not the definition of the platform.
Separate what changes at different speeds
Why did a small release sometimes demand a large review? Application source and environment selection were changing in the same place. The useful move was to separate three things that change for different reasons.
| Plane | Owns | Changes when |
|---|---|---|
| Infrastructure intent | Network, compute, identity, shared services | Platform requirements change |
| Application artifact | Tested executable or container | Application code changes |
| Deployment state | Environment-selected artifact and configuration | Release promotion or rollback |
Terraform declares cloud infrastructure, OIDC-based CI identity avoids long-lived build credentials, and managed builds publish immutable artifacts. A dedicated state repository records what each environment should run, and Argo CD reconciles Kubernetes workloads toward that reviewed intent. The review questions become narrower: “What does the software do?” belongs to the application change; “What should this environment run?” belongs to deployment state. Promotion copies a tested choice forward, and rollback selects a previous known artifact.
Kubernetes was a portfolio decision
Why operate Kubernetes at all when AWS already offers simpler hosting products? Lambda remains useful for bounded control work within its 15-minute limit. ECS with Fargate is viable for isolated long-running containers. A managed AI platform is attractive for a standard train-and-host loop. But FORGE also had to deliver custom pipelines, a Go decision service, experiment integration, edge artifacts, and conventional applications, and a separate ML-only lifecycle would have duplicated promotion, identity, and observability contracts.
EKS became reasonable because scheduled jobs, always-on services, specialized placement, GitOps, and observability could amortize one operating model. Its control-plane fee, worker costs, and upgrade burden remain real, so the rule is intentionally reversible: use Kubernetes while portfolio reuse exceeds its operating cost; choose a managed runtime whenever it does not. Processor choice follows the same rule. Compatible edge workloads run on Arm64, and cost efficiency is measured as effective compute cost divided by completed work.
Autoscaling has two control loops
Adding a node does not decide how many application replicas should exist. There are two control loops, and only one is automated today. Scheduled orchestration creates bounded pods with task-specific resource requests, while reviewed deployment state sets replica intent for long-running services. The infrastructure loop adds capacity when desired units cannot be placed. Fixed replica intent is defensible for the current serving workload because an availability floor and measured memory-sensitive capacity dominate a compact demand envelope. If request variability, queueing, freshness, or tail latency begins to consume that margin, the next controller should scale on the signal that represents work; CPU alone is often weak for a memory-bound service.
Three failures became platform contracts
The platform became more coherent when failures stopped being isolated fixes and started changing shared contracts.
- Memory pressure. Encoding, evaluation, ranking, and orchestration each reached peak memory differently. Streaming large data, chunking evaluation, reusing cached encodings, and recording resident-memory checkpoints made both failures and costs explainable. The lesson was task-aware sizing, not a larger default pod.
- Network-address pressure. A scale-up once looked successful from the compute side while placement was still blocked by address availability. Address capacity now sits beside CPU, memory, and specialized hardware in capacity reviews, and diagnostics distinguish “instance launched” from “pod network ready.”
- Ambiguous job outcomes. If success is inferred from output volume, a valid no-work run looks like a broken job. Pipeline contracts now distinguish success, skip, and failure, so deliberate no-work is reported without a false alarm while missing required input still fails closed.
Together they changed the contract: size by task, model addresses explicitly, and make terminal state semantic.
An alert is part of the system interface
A metric becomes operational when it reaches an owner with enough context to decide what to do next. FORGE uses Prometheus-compatible collection, managed metric storage, Grafana, alert policy, and Slack notification. Workloads define meaning; the platform supplies collection, retention, dashboards, and delivery. An actionable notification identifies impact, ownership, diagnostic context, and the next action, and rules carry volume guards, time windows, and resolved behavior.
One small build-versus-buy decision shows the cost discipline. The alert path uses a small Lambda formatter between SNS and Slack rather than an always-on notification service, with 256 MB of memory and a 10-second failure ceiling. A conservative ceiling of 1,000 notifications each running the full timeout would cost about $0.0417 of compute. Against a fixed $25 per month alert-to-Slack subscription, Lambda compute alone would need roughly 600,000 fully billed, timeout-length invocations to break even. Alertmanager grouping, inhibition, and bounded log retention keep both spend and alert fatigue down.
One platform, different runtime shapes
Recsys learning is scheduled, artifact-producing ML that reuses reproducible orchestration and resource isolation. Recsys serving is always-on and memory-sensitive, reusing controlled rollout, health, scaling, and monitoring. Presort, the computer-vision workload, adds an edge artifact lifecycle and remote health. Translation is a conventional managed service that reuses delivery standards without joining Kubernetes. Experiment tracking and data annotation are optional workspaces whose interactive services can be disabled between campaigns. The test that matters: a new shape should reuse the contracts and introduce only the runtime machinery its workload requires.
Keep the platform reversible
The next improvements are application-aware scaling, better capacity forecasting, less environment-specific deployment logic, and stronger automated deployment verification. Managed runtimes should remain preferable whenever Kubernetes adds no material value. FORGE succeeds when the next workload is cheaper to reason about, not merely easier to place on EKS, and when its most complex runtime remains optional.
The full write-up is at xuhuizhan5.github.io/project/forge. The most detailed consumer of the platform is Treverse Recsys, covered in Lot 003.
