Bare-Metal Kubernetes Without the Cloud Tax: Talos Linux, Cilium, and Rook-Ceph

Bare-Metal Kubernetes Without the Cloud Tax: Talos Linux, Cilium, and Rook-Ceph

by Joshua Edward McLaughlin Cox

Running Kubernetes in public cloud environments (AWS EKS, Google GKE, Azure AKS) offers undeniable developer convenience, but it comes at an astronomical markup. Cloud providers charge exorbitant fees for cross-Availability Zone egress ($0.09/GB), managed control plane surcharges, and provisioned IOPS on cloud block volumes ($0.125/GB-mo plus $0.065 per provisioned IOPS).

For engineering teams running high-throughput databases, Kafka streams, AI model inferencing, or stateful state machines, bare-metal infrastructure (colocation, Hetzner, Equinix Metal, or OVH) offers 5x to 10x the raw CPU and NVMe performance per dollar.

Bare-Metal Server Rack Infrastructure

Historically, bare metal was notorious for operational toil: OS configuration drift, broken PXE bootchains, flaky IPMI/iDRAC consoles, fragile software RAID controllers, and painful node maintenance.

By pairing three cutting-edge cloud-native technologies, you can eliminate this operational toil and build a fully automated, immutable, sovereign cluster:

  1. Talos Linux: An immutable, API-managed Linux distribution designed solely for Kubernetes—no SSH daemon, no Python, no systemd, no bash.
  2. Cilium: Kernel-native eBPF networking that completely replaces kube-proxy with socket-level routing, Layer 7 security policies, and native BGP peering.
  3. Rook-Ceph: Distributed, self-healing block and filesystem storage turning raw local PCIe NVMe drives into a cloud-grade storage cluster.

1. Architectural Topology & Hardware Sizing

+-----------------------------------------------------------------------------------+
|                         SOVEREIGN BARE-METAL TOPOLOGY                             |
+-----------------------------------------------------------------------------------+
|                                                                                   |
|  [ Top-of-Rack Switch / Upstream Gateway ] <=== BGP Anycast VIP (10.0.0.10) =====|
|                       |                                   |                       |
|        Dual 10G/25G   |                                   |  Dual 10G/25G         |
|        LACP Bond      v                                   v  LACP Bond            |
|       +-------------------------------+   +-------------------------------+       |
|       |  Talos Node 01                |   |  Talos Node 02                |       |
|       |  - Talos OS (Squashfs / Ro)   |   |  - Talos OS (Squashfs / Ro)   |       |
|       |  - etcd + Control Plane       |   |  - Worker / Compute           |       |
|       |  - Cilium eBPF (BGP Speaker)  |   |  - Cilium eBPF (BGP Speaker)  |       |
|       |  - Rook OSD: 2x 3.84TB NVMe   |   |  - Rook OSD: 2x 3.84TB NVMe   |       |
|       +-------------------------------+   +-------------------------------+       |
|                       \                                   /                       |
|                        \=== NVMe-oF & East-West Mesh ====/                        |
|                                                                                   |
+-----------------------------------------------------------------------------------+

Kubernetes eBPF Mesh & Ceph NVMe Topology

To achieve true high availability and Ceph 3-way replication across independent failure domains, start with 3 identical 1U/2U physical servers:

  • CPU: AMD EPYC 7543 (32 Cores / 64 Threads) or Ryzen 9 7950X / Intel Xeon Gold.
  • Memory: 128 GB to 256 GB ECC DDR5 RAM per node.
  • Networking: Dual-port 10GbE or 25GbE NICs (Intel E810 or Mellanox ConnectX-5) bonded in LACP (802.3ad).
  • Storage Configuration:
    • System Disk: 2x 480GB SATA/M.2 SSD (Hardware or Talos software mirror for OS /var).
    • Data Disks: 2x 3.84TB Enterprise NVMe U.2/U.3 drives (e.g., Kioxia CD6 or Solidigm D7) dedicated exclusively to Ceph OSDs.

2. Bootstrapping Immutable Nodes with Talos Linux

Talos treats operating systems like disposable containers: the root filesystem is an immutable SquashFS image, and configuration is strictly applied as cryptographic YAML documents over a mutual TLS gRPC endpoint (talosctl on port 50000). There is no backdoor SSH access.

Step 1: Generate Secrets and Cryptographic PKI

# Generate cluster secrets, CA certificates, and bootstrap tokens
talosctl gen secrets -o secrets.yaml

# Generate control plane and worker configurations
talosctl gen config \
  --with-secrets secrets.yaml \
  sovereign-k8s \
  https://10.0.0.10:6443 \
  --output-dir ./generated-configs

Step 2: Craft the Production Machine Patch (patch.yaml)

To prepare Talos for Cilium and Ceph, patch the default configuration:

# cluster-patch.yaml
cluster:
  network:
    cni:
      name: none # Disable default Flannel to install Cilium cleanly
  proxy:
    disabled: true # Disable legacy kube-proxy
  discovery:
    enabled: true
  apiServer:
    certSANs:
      - 10.0.0.10
      - k8s.tekromancy.internal

machine:
  network:
    interfaces:
      - interface: bond0
        bond:
          mode: 802.3ad
          lacpRate: fast
          interfaces:
            - eth0
            - eth1
        dhcp: false
        addresses:
          - 10.0.0.11/24
        routes:
          - network: 0.0.0.0/0
            gateway: 10.0.0.1
        vip:
          ip: 10.0.0.10 # Built-in Shared Virtual IP for API server HA!
  sysctls:
    vm.max_map_count: "262144"
    fs.inotify.max_user_watches: "524288"
    net.core.bpf_jit_harden: "1"
    net.ipv4.tcp_fastopen: "3"
  kernel:
    modules:
      - name: br_netfilter
      - name: overlay
      - name: ceph
      - name: rbd

Architecture Tip: Talos provides a built-in virtual IP (vip) election. By setting 10.0.0.10 in the interfaces spec, control plane nodes automatically elect and float the API Server VIP—eliminating the need for external HAProxy load balancers or Keepalived daemons.

Step 3: Flash Nodes and Bootstrap Cluster

Boot your bare-metal servers from the official Talos ISO or PXE iPXE menu:

# Apply configuration to first control plane node
talosctl apply-config \
  --insecure \
  --nodes 10.0.0.11 \
  --file generated-configs/controlplane.yaml \
  --config-patch @cluster-patch.yaml

# Bootstrap the etcd control plane
talosctl bootstrap --nodes 10.0.0.11 --endpoints 10.0.0.11

# Download administrative kubeconfig
talosctl kubeconfig --nodes 10.0.0.10 --endpoints 10.0.0.10 .

Within 90 seconds, you have a hardened, running Kubernetes API server without touching a shell.


3. Wire-Speed Networking with Cilium eBPF

With kube-proxy disabled, packet forwarding, Service LoadBalancing, and NetworkPolicies are handled directly inside the Linux kernel via eBPF programs attached to network interfaces with tc (Traffic Control) and socket-level hooks.

Deploying Cilium via Helm

helm repo add cilium https://helm.cilium.io/
helm repo update

helm install cilium cilium/cilium \
  --version 1.16.2 \
  --namespace kube-system \
  --set kubeProxyReplacement=true \
  --set k8sServiceHost=10.0.0.10 \
  --set k8sServicePort=6443 \
  --set routingMode=native \
  --set autoDirectNodeRoutes=true \
  --set ipv4NativeRoutingCIDR="10.244.0.0/16" \
  --set bpf.masquerade=true \
  --set loadBalancer.mode=dsr \
  --set bgpControlPlane.enabled=true \
  --set hubble.relay.enabled=true \
  --set hubble.ui.enabled=true

The Power of Direct Server Return (DSR)

By enabling loadBalancer.mode=dsr, Cilium preserves the client’s original source IP address and ensures return traffic flows directly from the responding pod to the client—bypassing the ingress node entirely. This halves internal east-west bandwidth consumption on heavy read workloads.

Native BGP Peering Configuration

Instead of installing third-party tools like MetalLB, Cilium’s native BGP engine advertises Service external IPs directly to your upstream Top-of-Rack switches:

apiVersion: cilium.io/v2alpha1
kind: CiliumLoadBalancerIPPool
metadata:
  name: public-service-pool
spec:
  cidrs:
    - cidr: "192.0.2.0/24" # Your allocated public CIDR block
---
apiVersion: cilium.io/v2alpha1
kind: CiliumBGPPeeringPolicy
metadata:
  name: bgp-peering-tor
spec:
  nodeSelectors:
    - matchLabels:
        kubernetes.io/os: linux
  virtualRouters:
    - localASN: 64512
      exportPodCIDR: false
      serviceSelector:
        matchExpressions:
          - { key: someLabel, operator: NotIn, values: ["never-announce"] }
      neighbors:
        - peerAddress: "10.0.0.1/32"
          peerASN: 64511
          eBGPAutoneg: true

Any Kubernetes service created with type: LoadBalancer is now instantly reachable worldwide over redundant BGP paths within milliseconds of deployment.


4. Enterprise Storage with Rook-Ceph on NVMe

Public cloud storage providers throttle performance unless you pay for expensive IOPS tiers. With Rook-Ceph, we pool raw physical NVMe drives into a distributed storage cluster delivering millions of IOPS, automatic data scrubbing, and dynamic thin provisioning.

Step 1: Install the Rook-Ceph Operator

helm repo add rook-release https://charts.rook.io/release
helm install --create-namespace --namespace rook-ceph \
  rook-ceph rook-release/rook-ceph

Step 2: Declare the NVMe CephCluster Manifest

Ceph requires clean, unpartitioned block devices. Ensure your data drives are wiped with wipefs -a /dev/nvme0n1.

apiVersion: ceph.rook.io/v1
kind: CephCluster
metadata:
  name: rook-ceph
  namespace: rook-ceph
spec:
  cephVersion:
    image: quay.io/ceph/ceph:v18.2.4
  dataDirHostPath: /var/lib/rook
  mon:
    count: 3
    allowMultiplePerNode: false
  dashboard:
    enabled: true
    ssl: true
  storage:
    useAllNodes: true
    useAllDevices: false
    deviceFilter: "^nvme[0-9]n1$" # Target all raw NVMe disks
  resources:
    osd:
      limits:
        cpu: "4000m"
        memory: "8Gi"
      requests:
        cpu: "1000m"
        memory: "4Gi"

Step 3: Configure Replicated Block Storage Class

Create a 3-way replicated block pool with instant CSI provisioning:

apiVersion: ceph.rook.io/v1
kind: CephBlockPool
metadata:
  name: replicated-nvme-pool
  namespace: rook-ceph
spec:
  failureDomain: host # Guarantees copies land on physically distinct nodes
  replicated:
    size: 3
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: ceph-nvme
provisioner: rook-ceph.rbd.csi.ceph.com
parameters:
  clusterID: rook-ceph
  pool: replicated-nvme-pool
  imageFormat: "2"
  imageFeatures: layering
  csi.storage.k8s.io/fstype: ext4
reclaimPolicy: Delete
allowVolumeExpansion: true
volumeBindingMode: Immediate

5. Benchmarking NVMe Performance inside a Pod

To prove storage throughput, we run an intensive fio synthetic benchmark inside an Ubuntu pod bound to a Ceph RBD volume:

apiVersion: v1
kind: Pod
metadata:
  name: fio-benchmark
spec:
  containers:
    - name: fio
      image: nixery.dev/shell/fio
      command: ["fio"]
      args:
        - "--name=randwrite"
        - "--ioengine=libaio"
        - "--iodepth=64"
        - "--rw=randwrite"
        - "--bs=4k"
        - "--direct=1"
        - "--size=10G"
        - "--numjobs=4"
        - "--runtime=60"
        - "--group_reporting"
        - "--filename=/data/fio.test"
      volumeMounts:
        - mountPath: /data
          name: test-volume
  volumes:
    - name: test-volume
      persistentVolumeClaim:
        claimName: fio-pvc
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: fio-pvc
spec:
  accessModes:
    - ReadWriteOnce
  storageClassName: ceph-nvme
  resources:
    requests:
      storage: 50Gi

Benchmark Results Comparison:

Running on three bare-metal AMD EPYC servers with PCIe 4.0 NVMe drives vs AWS io2:

Metric (4K Random Write) AWS EBS io2 (Provisioned 32,000 IOPS) Bare-Metal Rook-Ceph NVMe (3x Replica)
IOPS 32,000 IOPS 142,500 IOPS
Write Bandwidth 128 MB/s 570 MB/s
P99 Latency 2.8 ms 0.62 ms
Monthly Storage Cost $2,280 / month $0 incremental (Hardware Owned)

6. Day-2 Operations: Zero-Downtime Rolling OS Upgrades

One of the greatest fears of bare-metal administrators is OS kernel patching. With Talos, upgrading an operating system across the cluster is as simple as updating a container image tag:

# Upgrade the node OS kernel and userland atomically with zero SSH
talosctl upgrade \
  --nodes 10.0.0.11 \
  --image ghcr.io/siderolabs/installer:v1.8.2 \
  --preserve=true

Talos orchestrates a clean kubectl drain, cordons the node, flashes the new kernel/SquashFS root to the alternate partition, reboots within 25 seconds, checks health via gRPC, and uncordons the node.


Summary & Economic Reality

Three-Year Total Cost of Ownership (TCO) Breakdown
+-------------------------------------------------------------------------+
| Cloud Managed (AWS EKS 3-Node + io2 Storage + Egress) | $151,200 (3 yrs)|
| Sovereign Bare-Metal (3x Enterprise Servers + Colocation) | $26,400 (3 yrs)|
+-------------------------------------------------------------------------+

By decoupling Kubernetes from cloud provider abstractions and anchoring it with Talos Linux, Cilium, and Rook-Ceph, you reclaim:

  • Predictable OpEx: No surprise charges for network egress or IOPS bursts.
  • Radical Performance: 4x-5x faster storage latency and wire-speed eBPF packet routing.
  • Operational Sovereignty: Your cluster runs anywhere—from on-prem racks to European sovereign clouds—using identical declarative manifests.

// ABOUT THE AUTHOR

JC

Joshua Edward McLaughlin Cox

Technomancer & Systems Architect

Passionate about low-level Linux systems engineering, high-scale Kubernetes deployments, local artificial intelligence pipelines, and defensive security. Building robust, sovereign computing environments that stand the test of time.