Senior Systems Takeaway:
  • The 128KB Default Bottleneck: ZFS defaults to a 128KB recordsize on datasets. While safe for general file shares, it imposes up to a 45% write amplification penalty on databases and chokes throughput on multi-terabyte media archives.
  • Media Library Optimization (1M Recordsize): For datasets storing large sequential media files (>100MB 4K video, disk images, backups), tuning recordsize to 1MB (zfs set recordsize=1M pool/dataset) improves sequential read throughput by up to 35% and increases ZFS compression efficiency.
  • Database & VM Volblocksize Alignment: For PostgreSQL and MySQL, matching recordsize to database page size (8KB for Postgres, 16KB for MySQL InnoDB) eliminates the catastrophic read-modify-write cycle that destroys random IOPS.

When sysadmins build a high-performance TrueNAS SCALE pool or Proxmox ZFS storage target, attention is universally lavished on hardware: enterprise SAS3 HBAs, Micron PCIe 4.0 NVMe drives, ECC DDR5 memory, and 10GbE network cards. Yet after investing thousands in top-tier silicon, storage benchmarks frequently disappoint—crippled by mediocre IOPS or sluggish streaming speeds.

In over 80% of storage bottlenecks, the root cause is not hardware failure or controller throttling; it is misconfigured ZFS Recordsize. ZFS dynamically allocates blocks up to a designated maximum ceiling, which defaults to 128KB. Applying a one-size-fits-all 128KB block size across databases, virtual machine zvols, and 80GB 4K remuxes severely degrades both IOPS and storage longevity.

What Is ZFS Recordsize and How Does It Affect Performance?

Direct Answer: ZFS recordsize is the maximum logical block size used for storing files in a dataset. Matching recordsize to your workload eliminates write amplification: large blocks (1MB) maximize sequential throughput and compression for media streaming, while small blocks (8KB–16KB) prevent write-amplification penalties for databases and VMs.

To grasp why recordsize dictates performance, consider how ZFS writes data. When an application modifies an existing file, ZFS is a Copy-on-Write (CoW) filesystem. It does not overwrite blocks in place; it writes a new block elsewhere on disk and updates the metadata tree. If a database makes a discrete 8KB modification to a 128KB ZFS block, ZFS must read the entire 128KB block into RAM, update the 8KB portion, and write an entire new 128KB block to disk—amplifying disk writes sixteenfold.

Conversely, when storing large sequential files like video or raw image archives, a 128KB block size fractures a 50GB file into hundreds of thousands of individual metadata pointers, bloating the Adaptive Replacement Cache (ARC) and degrading sequential transfer rates.

Workload Tuning Matrix: Optimal ZFS Recordsize by Use Case

The configuration matrix below provides validated recordsize values, compression pairings, and expected performance gains across common homelab and enterprise workloads:

Workload Type Optimal Recordsize / Volblocksize Recommended Compression Performance Impact
4K Video & Plex/Jellyfin Media 1M (1024 KB) zstd-1 or lz4 30–35% faster sequential read, lower ARC metadata drag
Proxmox Backup / Borg / Restic 1M (1024 KB) zstd-3 Maximum compression ratio, accelerated deduplication chunking
General Office Documents / SMB Shares 128 KB (Default) lz4 Balanced random and sequential throughput for mixed files
PostgreSQL Database Datasets 8 KB lz4 Zero write amplification, 1:1 page alignment, 4x random IOPS
MySQL / MariaDB InnoDB Datasets 16 KB lz4 Direct match to InnoDB 16KB pages, eliminates read-modify-write
KVM Virtual Machine Disks (ZVOLs) 16 KB or 64 KB lz4 Eliminates ZFS pool fragmentation on NTFS/EXT4 guest writes

CLI Execution: Setting Recordsize and Rewriting Existing Data

An essential characteristic of ZFS recordsize tuning is that recordsize changes are not retroactive. Altering the property on an existing dataset only affects newly written data. To apply a new recordsize to existing files, data must be rewritten to disk:

# 1. Update dataset recordsize to 1MB for media storage
zfs set recordsize=1M tank/media

# 2. Verify dataset properties
zfs get recordsize tank/media

# 3. Rewrite existing data in-place to apply the new 1M block layout
find /mnt/tank/media -type f -exec cp --attributes-only --preserve=all {} {}.tmp \; -exec mv {}.tmp {} \;

For workloads requiring accelerated random access and metadata processing across spinning disk pools, combine recordsize tuning with a dedicated flash vdev as analyzed in our guide on ZFS Special vdevs for TrueNAS & Proxmox in 2026.

Senior Analyst’s Verdict: Never run an entire TrueNAS pool on the default 128KB recordsize. Structure your pool into purpose-driven datasets: set media libraries and backup archives to 1MB to unlock blazing sequential read speeds and compression, while isolating databases and VM virtual disks onto 8KB or 16KB datasets to eliminate write amplification. Tuning recordsize is a zero-cost configuration change that yields double-digit performance gains.

People Also Ask

Does setting recordsize to 1MB waste storage space for small files?
No. ZFS dynamic block allocation automatically scales block sizes down for files smaller than the recordsize ceiling. A 4KB text file stored on a 1MB dataset consumes only 4KB of disk space plus metadata, not 1MB.

Can I change the recordsize on a root ZFS pool without breaking TrueNAS?
Do not modify the recordsize of the boot pool or root dataset. Always leave system datasets at defaults and apply custom recordsize settings exclusively to child storage datasets created for user data.

What is the difference between recordsize and volblocksize in ZFS?
Recordsize applies to file-based ZFS datasets (POSIX filesystems). Volblocksize applies to block-level ZFS volumes (ZVOLs) used as virtual raw disks for Proxmox KVM virtual machines or iSCSI LUNs.