Platform Engineering

The platform your engineers build on, designed and run so they spend their time shipping product instead of fighting the infrastructure under it.

Who it is for
Teams whose own infrastructure has become the thing slowing them down, usually around the point where three or more teams deploy independently with different tooling.
What you get
  • Kubernetes on EKS or GKE
  • Terraform, modules and GitOps
  • Observability and DORA metrics
  • Backstage developer portal
  • Internal CLI tooling
  • Security hardening
What I need from you
Cloud account access, the current state of the infrastructure, and one engineer who knows it.
Overview

What this is

A platform is worth building when the same setup work happens in every team, in a slightly different way each time. Before that point it is overhead. After it, the cost of not having one shows up as a new service taking days to stand up, four teams each running their own CI, and nobody able to say how often anything is deployed.

I build the platform and then hand it over. That means infrastructure as code that reads like application code, a portal that answers "how do I start a new service" without a human in the loop, and observability that pages the person who can fix the thing rather than everyone.

Capabilities

What you get

Kubernetes on EKS and GKE

  • Cluster design: node groups, autoscaling, and scale-to-zero for burst workloads
  • Multi-tenancy with namespaces, quotas and network policy
  • GitOps delivery with ArgoCD, including app-of-apps and progressive sync
  • Upgrade paths that do not need a maintenance window

Terraform and infrastructure as code

  • Custom provider development, including generating one from gRPC protos
  • Reusable module libraries with semantic versioning and docs
  • Remote state with locking and encryption
  • Drift detection and automated remediation
  • Atlantis or Spacelift for plan-and-apply governance in pull requests

Observability and DORA metrics

  • Prometheus scrape config, recording rules and alerting rules
  • Grafana dashboards at the infrastructure, application and DORA layers
  • Loki for structured log aggregation
  • Alertmanager routing to PagerDuty, Slack or OpsGenie
  • SLO error budgets and burn-rate alerts per service

Added when an engagement needs it

Developer portals with Backstage

  • Deployment, configuration and a theme that matches your own
  • Software catalog ingested from GitHub, GitLab and Jira
  • Service templates that scaffold a new service with CI/CD already wired
  • Custom plugins for internal systems, DORA dashboards and on-call status
  • A tech radar, so standards are written down rather than remembered

Internal tooling and onboarding

  • Go CLIs with Cobra for environment setup, scaffolding and deployment
  • Interactive TUI workflows with Bubbletea for multi-step operations
  • Self-service provisioning: cloud credentials, Kubernetes access, VPN
  • Local development setup with devcontainers or Nix flakes
  • Distribution through Homebrew, binaries or Docker

Security hardening

  • Image scanning with Trivy and Snyk, in the pipeline rather than after it
  • Secret rotation and HashiCorp Vault integration
  • Network policy enforcement and pod security standards
  • Least-privilege IAM audits on AWS and GCP
  • Evidence gathering for SOC 2 and ISO 27001 readiness
Evidence

What I can show you

ICF is early, so there are no named client logos here yet. What there is: public code you can open and read, and work described without naming whose it was.

Approach

How I work

01

Discovery

Two weeks of reading what runs today and talking to the engineers who touch it. The output is a written picture of where time goes, which is usually not where people expect.

02

Design

A platform design you can argue with: what gets built, what gets bought, what stays as it is, and what each choice costs to run.

03

Build and iterate

Delivered in slices that are usable on their own. The first one is in your engineers hands before the second is designed.

04

Hand over

Runbooks, a walkthrough with the people who will own it, and a period where I am still around while they run it.

Technology

Stack

KubernetesAWS EKSGCP GKETerraformArgoCDBackstagePrometheusGrafanaLokiVaultGoHelm
When to engage

Signs this is the work you need

Your engineers manage Kubernetes instead of shipping product

Standing up a new service takes days rather than minutes

Three or more teams deploy independently, each with their own tooling

Nobody can tell you your deployment frequency or your time to restore

Related Services

You may also need

Is this the work you need?

Tell me what you are running and what is slowing you down. I reply within one business day, and I will say so if this is not work I should take.

Taking a small number of contracts