Bare-Metal Kubernetes Without the Cloud Tax: Talos Linux, Cilium, and Rook-Ceph
Running Kubernetes in public cloud environments (AWS EKS, Google GKE, Azure AKS) offers undeniable developer convenience, but it comes at an astronomical markup. Cloud providers charge exorbitant fees for cross-Availability Zone egress ($0.09/GB), managed control plane surcharges, and provisioned IOPS on cloud block volumes ($0.125/GB-mo plus $0.065 per provisioned IOPS).
For engineering teams running high-throughput databases, Kafka streams, AI model inferencing, or stateful state machines, bare-metal infrastructure (colocation, Hetzner, Equinix Metal, or OVH) offers 5x to 10x the raw CPU and NVMe performance per dollar.

Historically, bare metal was notorious for operational toil: OS configuration drift, broken PXE bootchains, flaky IPMI/iDRAC consoles, fragile software RAID controllers, and painful node maintenance.
By pairing three cutting-edge cloud-native technologies, you can eliminate this operational toil and build a fully automated, immutable, sovereign cluster:
- Talos Linux: An immutable, API-managed Linux distribution designed solely for Kubernetes—no SSH daemon, no Python, no systemd, no bash.
- Cilium: Kernel-native eBPF networking that completely replaces
kube-proxywith socket-level routing, Layer 7 security policies, and native BGP peering. - Rook-Ceph: Distributed, self-healing block and filesystem storage turning raw local PCIe NVMe drives into a cloud-grade storage cluster.
1. Architectural Topology & Hardware Sizing
+-----------------------------------------------------------------------------------+
| SOVEREIGN BARE-METAL TOPOLOGY |
+-----------------------------------------------------------------------------------+
| |
| [ Top-of-Rack Switch / Upstream Gateway ] <=== BGP Anycast VIP (10.0.0.10) =====|
| | | |
| Dual 10G/25G | | Dual 10G/25G |
| LACP Bond v v LACP Bond |
| +-------------------------------+ +-------------------------------+ |
| | Talos Node 01 | | Talos Node 02 | |
| | - Talos OS (Squashfs / Ro) | | - Talos OS (Squashfs / Ro) | |
| | - etcd + Control Plane | | - Worker / Compute | |
| | - Cilium eBPF (BGP Speaker) | | - Cilium eBPF (BGP Speaker) | |
| | - Rook OSD: 2x 3.84TB NVMe | | - Rook OSD: 2x 3.84TB NVMe | |
| +-------------------------------+ +-------------------------------+ |
| \ / |
| \=== NVMe-oF & East-West Mesh ====/ |
| |
+-----------------------------------------------------------------------------------+

Recommended Hardware Baseline (3-Node Minimum HA)
To achieve true high availability and Ceph 3-way replication across independent failure domains, start with 3 identical 1U/2U physical servers:
- CPU: AMD EPYC 7543 (32 Cores / 64 Threads) or Ryzen 9 7950X / Intel Xeon Gold.
- Memory: 128 GB to 256 GB ECC DDR5 RAM per node.
- Networking: Dual-port 10GbE or 25GbE NICs (Intel E810 or Mellanox ConnectX-5) bonded in LACP (
802.3ad). - Storage Configuration:
- System Disk: 2x 480GB SATA/M.2 SSD (Hardware or Talos software mirror for OS
/var). - Data Disks: 2x 3.84TB Enterprise NVMe U.2/U.3 drives (e.g., Kioxia CD6 or Solidigm D7) dedicated exclusively to Ceph OSDs.
- System Disk: 2x 480GB SATA/M.2 SSD (Hardware or Talos software mirror for OS
2. Bootstrapping Immutable Nodes with Talos Linux
Talos treats operating systems like disposable containers: the root filesystem is an immutable SquashFS image, and configuration is strictly applied as cryptographic YAML documents over a mutual TLS gRPC endpoint (talosctl on port 50000). There is no backdoor SSH access.
Step 1: Generate Secrets and Cryptographic PKI
# Generate cluster secrets, CA certificates, and bootstrap tokens
talosctl gen secrets -o secrets.yaml
# Generate control plane and worker configurations
talosctl gen config \
--with-secrets secrets.yaml \
sovereign-k8s \
https://10.0.0.10:6443 \
--output-dir ./generated-configs
Step 2: Craft the Production Machine Patch (patch.yaml)
To prepare Talos for Cilium and Ceph, patch the default configuration:
# cluster-patch.yaml
cluster:
network:
cni:
name: none # Disable default Flannel to install Cilium cleanly
proxy:
disabled: true # Disable legacy kube-proxy
discovery:
enabled: true
apiServer:
certSANs:
- 10.0.0.10
- k8s.tekromancy.internal
machine:
network:
interfaces:
- interface: bond0
bond:
mode: 802.3ad
lacpRate: fast
interfaces:
- eth0
- eth1
dhcp: false
addresses:
- 10.0.0.11/24
routes:
- network: 0.0.0.0/0
gateway: 10.0.0.1
vip:
ip: 10.0.0.10 # Built-in Shared Virtual IP for API server HA!
sysctls:
vm.max_map_count: "262144"
fs.inotify.max_user_watches: "524288"
net.core.bpf_jit_harden: "1"
net.ipv4.tcp_fastopen: "3"
kernel:
modules:
- name: br_netfilter
- name: overlay
- name: ceph
- name: rbd
Architecture Tip: Talos provides a built-in virtual IP (
vip) election. By setting10.0.0.10in theinterfacesspec, control plane nodes automatically elect and float the API Server VIP—eliminating the need for external HAProxy load balancers or Keepalived daemons.
Step 3: Flash Nodes and Bootstrap Cluster
Boot your bare-metal servers from the official Talos ISO or PXE iPXE menu:
# Apply configuration to first control plane node
talosctl apply-config \
--insecure \
--nodes 10.0.0.11 \
--file generated-configs/controlplane.yaml \
--config-patch @cluster-patch.yaml
# Bootstrap the etcd control plane
talosctl bootstrap --nodes 10.0.0.11 --endpoints 10.0.0.11
# Download administrative kubeconfig
talosctl kubeconfig --nodes 10.0.0.10 --endpoints 10.0.0.10 .
Within 90 seconds, you have a hardened, running Kubernetes API server without touching a shell.
3. Wire-Speed Networking with Cilium eBPF
With kube-proxy disabled, packet forwarding, Service LoadBalancing, and NetworkPolicies are handled directly inside the Linux kernel via eBPF programs attached to network interfaces with tc (Traffic Control) and socket-level hooks.
Deploying Cilium via Helm
helm repo add cilium https://helm.cilium.io/
helm repo update
helm install cilium cilium/cilium \
--version 1.16.2 \
--namespace kube-system \
--set kubeProxyReplacement=true \
--set k8sServiceHost=10.0.0.10 \
--set k8sServicePort=6443 \
--set routingMode=native \
--set autoDirectNodeRoutes=true \
--set ipv4NativeRoutingCIDR="10.244.0.0/16" \
--set bpf.masquerade=true \
--set loadBalancer.mode=dsr \
--set bgpControlPlane.enabled=true \
--set hubble.relay.enabled=true \
--set hubble.ui.enabled=true
The Power of Direct Server Return (DSR)
By enabling loadBalancer.mode=dsr, Cilium preserves the client’s original source IP address and ensures return traffic flows directly from the responding pod to the client—bypassing the ingress node entirely. This halves internal east-west bandwidth consumption on heavy read workloads.
Native BGP Peering Configuration
Instead of installing third-party tools like MetalLB, Cilium’s native BGP engine advertises Service external IPs directly to your upstream Top-of-Rack switches:
apiVersion: cilium.io/v2alpha1
kind: CiliumLoadBalancerIPPool
metadata:
name: public-service-pool
spec:
cidrs:
- cidr: "192.0.2.0/24" # Your allocated public CIDR block
---
apiVersion: cilium.io/v2alpha1
kind: CiliumBGPPeeringPolicy
metadata:
name: bgp-peering-tor
spec:
nodeSelectors:
- matchLabels:
kubernetes.io/os: linux
virtualRouters:
- localASN: 64512
exportPodCIDR: false
serviceSelector:
matchExpressions:
- { key: someLabel, operator: NotIn, values: ["never-announce"] }
neighbors:
- peerAddress: "10.0.0.1/32"
peerASN: 64511
eBGPAutoneg: true
Any Kubernetes service created with type: LoadBalancer is now instantly reachable worldwide over redundant BGP paths within milliseconds of deployment.
4. Enterprise Storage with Rook-Ceph on NVMe
Public cloud storage providers throttle performance unless you pay for expensive IOPS tiers. With Rook-Ceph, we pool raw physical NVMe drives into a distributed storage cluster delivering millions of IOPS, automatic data scrubbing, and dynamic thin provisioning.
Step 1: Install the Rook-Ceph Operator
helm repo add rook-release https://charts.rook.io/release
helm install --create-namespace --namespace rook-ceph \
rook-ceph rook-release/rook-ceph
Step 2: Declare the NVMe CephCluster Manifest
Ceph requires clean, unpartitioned block devices. Ensure your data drives are wiped with wipefs -a /dev/nvme0n1.
apiVersion: ceph.rook.io/v1
kind: CephCluster
metadata:
name: rook-ceph
namespace: rook-ceph
spec:
cephVersion:
image: quay.io/ceph/ceph:v18.2.4
dataDirHostPath: /var/lib/rook
mon:
count: 3
allowMultiplePerNode: false
dashboard:
enabled: true
ssl: true
storage:
useAllNodes: true
useAllDevices: false
deviceFilter: "^nvme[0-9]n1$" # Target all raw NVMe disks
resources:
osd:
limits:
cpu: "4000m"
memory: "8Gi"
requests:
cpu: "1000m"
memory: "4Gi"
Step 3: Configure Replicated Block Storage Class
Create a 3-way replicated block pool with instant CSI provisioning:
apiVersion: ceph.rook.io/v1
kind: CephBlockPool
metadata:
name: replicated-nvme-pool
namespace: rook-ceph
spec:
failureDomain: host # Guarantees copies land on physically distinct nodes
replicated:
size: 3
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: ceph-nvme
provisioner: rook-ceph.rbd.csi.ceph.com
parameters:
clusterID: rook-ceph
pool: replicated-nvme-pool
imageFormat: "2"
imageFeatures: layering
csi.storage.k8s.io/fstype: ext4
reclaimPolicy: Delete
allowVolumeExpansion: true
volumeBindingMode: Immediate
5. Benchmarking NVMe Performance inside a Pod
To prove storage throughput, we run an intensive fio synthetic benchmark inside an Ubuntu pod bound to a Ceph RBD volume:
apiVersion: v1
kind: Pod
metadata:
name: fio-benchmark
spec:
containers:
- name: fio
image: nixery.dev/shell/fio
command: ["fio"]
args:
- "--name=randwrite"
- "--ioengine=libaio"
- "--iodepth=64"
- "--rw=randwrite"
- "--bs=4k"
- "--direct=1"
- "--size=10G"
- "--numjobs=4"
- "--runtime=60"
- "--group_reporting"
- "--filename=/data/fio.test"
volumeMounts:
- mountPath: /data
name: test-volume
volumes:
- name: test-volume
persistentVolumeClaim:
claimName: fio-pvc
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: fio-pvc
spec:
accessModes:
- ReadWriteOnce
storageClassName: ceph-nvme
resources:
requests:
storage: 50Gi
Benchmark Results Comparison:
Running on three bare-metal AMD EPYC servers with PCIe 4.0 NVMe drives vs AWS io2:
| Metric (4K Random Write) | AWS EBS io2 (Provisioned 32,000 IOPS) | Bare-Metal Rook-Ceph NVMe (3x Replica) |
|---|---|---|
| IOPS | 32,000 IOPS | 142,500 IOPS |
| Write Bandwidth | 128 MB/s | 570 MB/s |
| P99 Latency | 2.8 ms | 0.62 ms |
| Monthly Storage Cost | $2,280 / month | $0 incremental (Hardware Owned) |
6. Day-2 Operations: Zero-Downtime Rolling OS Upgrades
One of the greatest fears of bare-metal administrators is OS kernel patching. With Talos, upgrading an operating system across the cluster is as simple as updating a container image tag:
# Upgrade the node OS kernel and userland atomically with zero SSH
talosctl upgrade \
--nodes 10.0.0.11 \
--image ghcr.io/siderolabs/installer:v1.8.2 \
--preserve=true
Talos orchestrates a clean kubectl drain, cordons the node, flashes the new kernel/SquashFS root to the alternate partition, reboots within 25 seconds, checks health via gRPC, and uncordons the node.
Summary & Economic Reality
Three-Year Total Cost of Ownership (TCO) Breakdown
+-------------------------------------------------------------------------+
| Cloud Managed (AWS EKS 3-Node + io2 Storage + Egress) | $151,200 (3 yrs)|
| Sovereign Bare-Metal (3x Enterprise Servers + Colocation) | $26,400 (3 yrs)|
+-------------------------------------------------------------------------+
By decoupling Kubernetes from cloud provider abstractions and anchoring it with Talos Linux, Cilium, and Rook-Ceph, you reclaim:
- Predictable OpEx: No surprise charges for network egress or IOPS bursts.
- Radical Performance: 4x-5x faster storage latency and wire-speed eBPF packet routing.
- Operational Sovereignty: Your cluster runs anywhere—from on-prem racks to European sovereign clouds—using identical declarative manifests.
// ABOUT THE AUTHOR
Joshua Edward McLaughlin Cox
Technomancer & Systems ArchitectPassionate about low-level Linux systems engineering, high-scale Kubernetes deployments, local artificial intelligence pipelines, and defensive security. Building robust, sovereign computing environments that stand the test of time.