Building a 3-node Proxmox VE Ceph cluster provides enterprise-grade High Availability (HA) and distributed block storage, allowing virtual machines to migrate live or fail over instantly if a physical node dies. By implementing a Full Mesh 10GbE routed network interconnect using dual-port SFP+ NICs, homelab engineers eliminate the single point of failure and cost of a 10GbE switch while providing Ceph with the low-latency bandwidth required for NVMe OSD replication.
In traditional homelab environments, storage redundancy is usually confined to a single box: a TrueNAS server with RAIDZ2 or a Proxmox host with mirrored ZFS NVMe pools. While this protects against drive failure, it represents a fatal single point of failure (SPOF)—if the host motherboard, power supply, or CPU dies, all hosted services go offline simultaneously.
Enter Ceph: the open-source, massively scalable distributed storage platform natively integrated into Proxmox VE. In 2026, with used enterprise 10GbE Mellanox ConnectX-3/4 NICs available for under $35 on eBay and consumer PCIe 4.0 NVMe drives offering sustained multi-gigabyte throughput, deploying a production-grade 3-node Ceph cluster at home has never been more accessible.
Can You Run Ceph on a 3-Node Proxmox Homelab?
Yes, 3 nodes is the exact minimum quorum requirement to run Ceph on Proxmox VE. With 3 nodes and a pool rule of size 3 / min_size 2, Ceph can tolerate the complete failure of an entire physical server without losing storage quorum, allowing Proxmox HA to automatically reboot virtual machines on the surviving nodes in under 60 seconds.
The mathematical heart of Ceph reliability is its CRUSH map and replication rules. In a standard 3-node cluster:
- Pool Size = 3: Every piece of data written by a VM is replicated across three distinct physical nodes.
- Min Size = 2: As long as at least two nodes agree on data integrity, Ceph continues accepting read and write operations. If one node loses power, the cluster enters a “degraded” state but remains 100% operational with zero data loss.
To understand the filesystem mechanics of local storage vs distributed block storage, review our benchmark analysis on Btrfs vs ZFS on Proxmox VE and RAM overhead.
The Switchless Full Mesh 10GbE Architecture
Ceph’s primary operational bottleneck is network latency. Because every write must be confirmed across multiple nodes before completing, running Ceph over a standard 1GbE network causes severe VM I/O wait lockups. However, buying a managed 10GbE SFP+ switch with redundant power supplies is noisy and expensive.
The elegant homelab solution is a Switchless Full Mesh Network:
| Interconnect Topology | Hardware Required | Network Latency | Single Point of Failure (SPOF) | Cost Profile |
|---|---|---|---|---|
| Full Mesh 10GbE (Routed) | 3x Dual-Port Mellanox NICs + 3 DAC cables | 0.04 ms (Direct P2P wire) | Zero (Fully redundant paths) | Under $120 total |
| 10G Managed Switch | 1x 10G Switch + 3x Single-Port NICs | 0.12 ms (Switch hop) | High (If switch dies, cluster splits) | $250 – $500 |
| 1GbE Standard LAN | Standard gigabit onboard NICs | 0.85 ms (High I/O wait) | Severe bottleneck | $0 (Unusable for Ceph) |
Configuring Full Mesh Routing via Debian Interfaces
In a 3-node full mesh, Node 1 connects directly to Node 2 and Node 3 using direct-attach copper (DAC) cables. Using static host routes in /etc/network/interfaces or Open vSwitch (OVS), each node communicates with peer nodes over direct point-to-point subnets (e.g., 10.10.12.0/24, 10.10.13.0/24, 10.10.23.0/24). If one cable fails, broadcast traffic automatically routes through the remaining peer.
Ceph NVMe OSD Tuning Rules for Homelab
When deploying modern M.2 NVMe solid-state drives as Ceph OSDs (Object Storage Daemons), follow these mandatory performance tuning rules:
- Never Use Consumer QLC SSDs: QLC drives exhaust their SLC write cache within seconds of sequential replication, causing Ceph write speeds to collapse from 1,500 MB/s to 60 MB/s, triggering OSD timeout heartbeats. Always use TLC NVMe drives with high sustained write endurance (such as Samsung 990 Pro, Kioxia Exceria Plus, or refurbished enterprise Kioxia/Micron drives).
- Separate Cluster and Public Traffic: Assign Ceph cluster traffic (peer-to-peer data replication) to the 10GbE full mesh network, while keeping the Ceph public network (VM disk I/O and Proxmox API) on a dedicated bridge to prevent replication spikes from starving VM access.
- Disable Write Cache Buffers on Low-Power Mini PCs: If running Proxmox on mini PCs (Intel N100/N305) without battery-backed write caches, configure
osd_memory_targetto 2GB per OSD to prevent Linux Out-of-Memory (OOM) killer crashes.
A 3-node Proxmox Ceph cluster connected via switchless full-mesh 10GbE delivers enterprise reliability for under $1,200 total hardware investment. While it consumes slightly more power than a single monolithic NAS, the ability to unplug the power cord from any random server in your rack while your VMs and Docker containers continue running uninterrupted is the holy grail of homelab engineering.
People Also Ask
Can I run Ceph with only 2 nodes in Proxmox?
Technically you can configure Ceph on 2 nodes with an external QDevice for quorum, but it is strongly discouraged for production storage. In a 2-node setup, if one node fails, the cluster cannot rebalance data, creating severe split-brain risks. Three nodes is the universally recommended minimum.
Is 10GbE required for Proxmox Ceph?
Yes, 10GbE (or faster) is practically mandatory for Ceph. Running Ceph over 1GbE causes severe network congestion during VM writes and OSD recovery, resulting in sluggish VM performance and high I/O wait timeouts.
How much usable storage do you get with a 3-node Ceph cluster?
Under the standard 3x replication policy (size 3), your usable capacity is exactly one-third (33.3%) of your total raw disk capacity across all nodes. For example, three 2TB NVMe drives (6TB raw) yield 2TB of usable high-availability storage.

