Skip to content
all systems normal Book assessment
// PLATFORM & ARCHITECTURE

One operating layer over the chaos you inherited.

The PsychoK control plane collapses 11 fragmented dashboards and four layers of stacked YAML into a single opinionated surface — ingest, reconcile, observe, and inject failure on the clusters you actually run in production.

Built in Berlin by three ex-CoreOS platform engineers. Running in 380+ production clusters at companies including Plaid, Mercury, Ramp, Vanta, and Linear.

  • 4.2×tooling spend cut, 90 days post-migration
  • 6 minaverage incident MTTR (industry: 47 min)
  • 99.99%uptime SLA, follow-the-sun SRE
// 01 — Architecture

Four planes. One opinionated surface. No YAML pyramids.

PsychoK runs four cooperating planes inside your cluster boundary. Each one replaces a category of vendor you would otherwise glue together with shell scripts and Slack approvals.

  1. 01

    Ingest plane

    A drop-in OTel collector fronted by a Prometheus-compatible remote_write endpoint. In 2024 we processed 9.4 trillion container metrics and 2.1 billion distributed traces without a single dropped sample at p99 — your dashboards stay accurate even when 200 nodes page at once.

    receivers:
      otlp:
        protocols: { grpc: {}, http: {} }
      prometheus:
        config:
          scrape_configs: [psychok-managed]
    processors:
      batch: { timeout: 1s, send_batch_size: 8192 }
    exporters:
      psychok/ingest: { region: eu-central-1 }
  2. 02

    Control plane

    GitOps reconciliation against a curated catalog of 2,100+ vetted Helm charts and ArgoCD ApplicationSets, refreshed every 6 hours. Drift is detected, diffed, and either auto-pruned or surfaced for review — whichever your policy dictates. We do not own your manifests; we own the loop.

    apiVersion: argoproj.io/v1alpha1
    kind: ApplicationSet
    metadata: { name: psychok-platform }
    spec:
      generators:
        - list: { elements: [ { cluster: prod-eu-1, release: platform-v2024.11 } ] }
      template: { releaseName: '{ {cluster} }-platform' }
  3. 03

    Observability plane

    Pre-aggregated SLO views per workload, plus a query layer that speaks PromQL, LogQL, and SQL against the same underlying columnar store. Median query returns in 380ms across a 14-day window — so the on-call engineer gets an answer before the page times out, not after.

    380ms median query, 14d window · 4.2× toolchain spend cut

  4. 04

    Chaos plane

    ChaosK, our in-house fault-injection engine, ran 14,000+ node-failure recoveries in customer clusters during 2024 alone — pod eviction, AZ failover, etcd quorum loss, kubelet restart storms. Exercises are scheduled, scoped to namespaces, and gated on SLO budget so chaos never eats the budget it is meant to protect.

    14,000+ node-failure recoveries in customer clusters, 2024

INGEST OTel · remote_write
CONTROL ArgoCD · Helm · OPA
OBSERVE SLOs · PromQL · LogQL
CHAOS ChaosK · game days
Schematic — the four planes run in-cluster and reconcile against a single external pilot cell.
// 02 — Integrations

Connective tissue for the stack you already run.

PsychoK is not a rip-and-replace. We sit beside every managed distribution and every GitOps tool your team has already vetted — extending, hardening, and unifying them under one operating layer.

/distros

Managed Kubernetes

First-class controllers for Amazon EKS, Google GKE, Azure AKS, Red Hat OpenShift, and vanilla kubeadm. Cluster adoption is a single CLI call; lifecycle patches, version skew, and node-pool rightsizing are handled by the control plane.

  • EKS — IRSA, Fargate profiles, Karpenter integration
  • GKE — Autopilot mode, workload identity, config connectors
  • AKS — Azure AD workload identity, Azure CNI overlay
  • OpenShift — Operator lifecycle, SCC translation
/gitops

GitOps & packaging

Native reconciliation via ArgoCD ApplicationSets, Flux v2 Kustomizations, and a curated catalog of 2,100+ vetted Helm charts refreshed every 6 hours. We auto-detect drift and file the PR — your reviewers approve, never author.

  • ArgoCD — ApplicationSet generator, sync windows
  • Flux v2 — Kustomization, HelmRelease, Image automation
  • Helm 3 — strict mode, OCI registries, values linting
  • Kustomize — overlay graph, secret decryption via ESO
/telemetry

Telemetry & traces

OpenTelemetry-native ingest, Prometheus-compatible remote_write, and a columnar store that holds 2.1B traces and 9.4T metrics without a single dropped sample at p99. Ship to your existing Datadog or New Relic if you want — or skip them and save the spend.

  • OpenTelemetry — OTLP gRPC & HTTP receivers
  • Prometheus — remote_write, exemplars, WAL replay
  • Jaeger / Tempo — trace export, tail sampling
  • Fluent Bit — log routing, redaction, PII scrubbing
/policy

Policy, secrets & identity

OPA Gatekeeper and Kyverno policies run at admission time; External Secrets Operator fans out to Vault, AWS Secrets Manager, and GCP Secret Manager. SPIFFE/SPIRE identities for service-to-service mTLS, audited end-to-end.

  • OPA Gatekeeper — Rego constraints, audit mode
  • Kyverno — mutation + validation, exception windows
  • External Secrets — Vault, AWS SM, GCP SM, Azure KV
  • SPIFFE / SPIRE — workload identities, mTLS rotation

Open-source maintainers of chaosk, argocd-precheck, and k8s-cost-explorer — 31,400+ combined GitHub stars, founding contributor to the CNCF-sandboxed OpenFeature K8s operator (March 2024).

// 02 · Capability matrix

What PsychoK owns — and what it deliberately doesn't.

Six capability lanes, all GA, all in production at customer scale. Maturity is reported honestly: tier-1 lanes are core; tier-2 lanes extend the platform; tier-3 lanes are evolving.

Capability In scope Out of scope Maturity Ref.
Cluster lifecycle EKS, GKE, AKS, self-managed K8s; provisioning, upgrades, node pools, scale-to-zero Bare-metal provisioning, OpenShift distribution, VMware Tanzu replacement Tier 1 docs/cluster-lifecycle
GitOps delivery ArgoCD ApplicationSets, 2,100+ vetted Helm charts, drift detection, pre-merge validation Flux CD migration tooling, Jenkins pipelines, Spinnaker Tier 1 docs/gitops
Observability Prometheus + Thanos, OpenTelemetry pipelines, 2.1B traces/yr processed, SLO burn alerts Log-only products, RUM/browser telemetry, custom dashboards-as-a-service Tier 1 docs/observability
Cost & FinOps Right-sizing, spot orchestration, showback/chargeback, $0.04/pod-hour cap AWS Cost Explorer replacement, custom invoice parsing Tier 1 docs/cost
Chaos & resilience In-house ChaosK engine, game-day scheduling, 14,000+ node failures recovered in 2024 Multi-cloud blast-radius drills, region-level failover Tier 2 docs/chaos
Security & compliance SOC 2 Type II, ISO 27001, HIPAA alignment, OPA policies, SBOM, signed releases FedRAMP authorization, SIEM replacement, custom GRC workflows Tier 1 docs/security

Self-qualify: if 4 of 6 lanes match your stack, a 30-minute assessment will confirm fit. Book assessment →

// 03 · Security & trust posture

The paperwork your security officer will ask for — already filed.

PsychoK is built for teams whose auditors read every line. Compliance is a deliverable, not a marketing badge.

Compliance

SOC 2 Type II

ISO 27001 · HIPAA-aligned

Audited annually. Reports available under NDA through the trust portal.

Penetration testing

3 firms

Cure53 · Trail of Bits · NCC Group

Consecutive third-party tests. Executive summaries published; full reports on request.

Customer retention

94% over 3 years

Industry benchmark: 71%

Measured across 1,840+ engineering teams on multi-year contracts.

Support SLA

4m 12s

Median first-response · 24/7 follow-the-sun

Staffed by ex-Google, ex-Spotify, and ex-HashiCorp engineers. 99.99% uptime SLA.

SOC 2 Type II ISO 27001 HIPAA-aligned CNCF member OpenFeature K8s operator author G2 Leader · Container Orchestration

// 04 · Next step

Thirty minutes with a senior engineer. No deck, no SDR.

Bring one production cluster, one pain point, one pager screenshot. We'll show you the control plane on your own topology and leave you with a written findings doc whether or not you buy.

Senior engineer on the call · Median response time 4 minutes 12 seconds · No procurement loop

  • Format30 min · live screen-share · your topology
  • OutcomeWritten findings + fit/no-fit call
  • Follow-upOptional 90-day migration plan, quoted fixed-fee