Module Parameters

Note

Most of this page is generated from the OpenZFS sources: the list of parameters, their types, defaults and descriptions come from the code and the man pages of each release. The tuning advice is written by hand in docs/module_parameters.yaml.

If anything here is wrong, outdated or missing, please report it.

Most OpenZFS kernel module parameters are accessible in the SysFS /sys/module/zfs/parameters directory. Current values can be observed by

cat /sys/module/zfs/parameters/PARAMETER

Many of these can be changed by writing new values. These are denoted by Change: Dynamic in the parameter details below.

echo NEWVALUE >> /sys/module/zfs/parameters/PARAMETER

If the parameter is not dynamically adjustable, an error can occur and the value will not be set. It can be helpful to check the permissions for the PARAMETER file in SysFS.

In some cases, the parameter must be set prior to loading the kernel modules or it is desired to have the parameters set automatically at boot time. For many distros, this can be accomplished by creating a file named /etc/modprobe.d/zfs.conf containing a text line for each module parameter using the format:

# change PARAMETER for workload XZY to solve problem PROBLEM_DESCRIPTION
# changed by YOUR_NAME on DATE
options zfs PARAMETER=VALUE

Some parameters related to ZFS operations are located in module parameters other than in the zfs kernel module. For example, the icp kernel module parameters are visible in the /sys/module/icp/parameters directory and can be set by default at boot time by changing the /etc/modprobe.d/icp.conf file.

See the man page for modprobe.d for more information.

On FreeBSD the same tunables are exposed as sysctls under the vfs.zfs tree, for example sysctl vfs.zfs.arc.max, and can be set at boot time from /boot/loader.conf.

To observe the list of parameters supported by the modules you actually have installed, along with a short synopsis of each, use the modinfo command:

modinfo zfs

Manual pages

The zfs(4) and spl(4) man pages (previously zfs- and spl-module-parameters(5), respectively, prior to OpenZFS 2.1) are the authoritative description of the module parameters and are shipped with the version of OpenZFS you run.

This page is generated from the same sources for every supported release, and adds the information that the man pages do not carry: which versions a parameter exists in, how its default changed between them, and the accumulated advice of ZFS developers and practitioners about when to change it.

Use the selector above to restrict the page to the parameters that exist in a particular OpenZFS release, and the box next to it to narrow the list down by name. Leave the selector at All versions to search the full history, including parameters that have since been removed.

Tags

The list of parameters is large and resists hierarchical representation. Each parameter is tagged with keywords for frequent searches.

ABD

allocation

ARC

BRT

btree

channel_programs

checkpoint

checksum

compression

condense

CPU

ctldir

dataset

dbuf

dbuf_cache

DDT

ddt_log

ddt_zap

deadman

debug

dedup

delay

delete

discard

disks

DMU

dmu_recv

dmu_send

dmu_traverse

dmu_zfetch

dnode

DSL

dsl_dataset

dsl_deadlist

dsl_pool

dsl_scan

encryption

file

filesystem

fm

fragmentation

gcm

generic

HDD

hostid

ICP

import

ioctl

kmem

kmem_cache

L2ARC

livelist

livelist_condense

lua

memory

metadata

metaslab

mirror

MMP

multihost

panic

prefetch

QAT

qat_compress

qat_crypt

raidz

receive

refcount

remove

resilver

scrub

send

snapshot

SPA

spa_config

spa_log_spacemap

spa_stats

special_vdev

SPL

SSD

super

taskq

trim

TXG

vdev

vdev_cache

vdev_disk

vdev_file

vdev_indirect

vdev_initialize

vdev_mirror

vdev_queue

vdev_rebuild

vdev_removal

vdev_trim

vfsops

vnops

volume

write_throttle

xattr

ZAP

zap_fat

zcp

zed

ZIL

ZIO

ZIO_scheduler

znode

zstd

ZVOL

Parameters

brt_zap_default_bs

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.4 - master

13

v2.2 - v2.3

12

Change:

Dynamic

Tags:

BRT

Default BRT ZAP data block size as a power of 2. Note that changing this after creating a BRT on the pool will not affect existing BRTs, only newly created ones.

brt_zap_default_ibs

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.4 - master

13

v2.2 - v2.3

12

Change:

Dynamic

Tags:

BRT

Default BRT ZAP indirect block size as a power of 2. Note that changing this after creating a BRT on the pool will not affect existing BRTs, only newly created ones.

brt_zap_prefetch

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

BRT, prefetch

Controls prefetching BRT records for blocks which are going to be cloned.

dbuf_cache_hiwater_pct

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Change:

Dynamic

Tags:

dbuf, dbuf_cache

The percentage over dbuf_cache_max_bytes when dbufs must be evicted directly.

When to change: Testing dbuf cache algorithms

Notes: The dbuf_cache_hiwater_pct and dbuf_cache_lowater_pct define the operating range for dbuf cache evict thread. The hiwater and lowater are percentages of the dbuf_cache_max_bytes value. When the dbuf cache grows above ((100% + dbuf_cache_hiwater_pct) * dbuf_cache_max_bytes) then the dbuf cache thread begins evicting. When the dbug cache falls below ((100% - dbuf_cache_lowater_pct) * dbuf_cache_max_bytes) then the dbuf cache thread stops evicting.

dbuf_cache_lowater_pct

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Change:

Dynamic

Tags:

dbuf, dbuf_cache

The percentage below dbuf_cache_max_bytes when the evict thread stops evicting dbufs.

When to change: Testing dbuf cache algorithms

Notes: The dbuf_cache_hiwater_pct and dbuf_cache_lowater_pct define the operating range for dbuf cache evict thread. The hiwater and lowater are percentages of the dbuf_cache_max_bytes value. When the dbuf cache grows above ((100% + dbuf_cache_hiwater_pct) * dbuf_cache_max_bytes) then the dbuf cache thread begins evicting. When the dbug cache falls below ((100% - dbuf_cache_lowater_pct) * dbuf_cache_max_bytes) then the dbuf cache thread stops evicting.

dbuf_cache_max_bytes

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.2 - master

UINT64_MAX

v2.0 - v2.1

ULONG_MAX

v0.8

0

Range:

0 = use dbuf_cache_shift to ARC c_max

Change:

Dynamic

Tags:

ARC, dbuf, dbuf_cache

Maximum size in bytes of the dbuf cache. The target size is determined by the MIN versus 1/2^dbuf_cache_shift (1/32nd) of the target ARC size. The behavior of the dbuf cache and its associated settings can be observed via the /proc/spl/kstat/zfs/dbufstats kstat.

Notes: dbuf_cache_max_bytes sets the size of the dbuf cache in bytes. This is an alternate method for setting dbuf cache size than dbuf_cache_shift Performance tuning of dbuf cache can be monitored using: - dbufstat command - node_exporter ZFS module for prometheus environments - telegraf ZFS plugin for general-purpose metric collection - /proc/spl/kstat/zfs/dbufstats kstat

dbuf_cache_max_shift

Versions:

v0.7

Platforms:

Linux, FreeBSD

Type:

int

Range:

1 to 63

Change:

Dynamic

Tags:

dbuf, dbuf_cache

Cap the size of the dbuf cache to a log2 fraction of arc size.

When to change: Testing dbuf cache algorithms

Notes: The dbuf_cache_max_bytes minimum is the lesser of dbuf_cache_max_bytes and the current ARC target size (c) >> dbuf_cache_max_shift

dbuf_cache_shift

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

5

Range:

5 to MAX_INT

Change:

Dynamic

Tags:

ARC, dbuf, dbuf_cache

Set the size of the dbuf cache (dbuf_cache_max_bytes) to a log2 fraction of the target ARC size.

When to change: to improve performance of read-intensive channel programs

Notes: dbuf_cache_shift sets the size of the dbuf cache as a fraction of ARC target size. This is an alternate method for setting dbuf cache size than dbuf_cache_max_bytes. dbuf_cache_max_bytes overrides dbuf_cache_shift This value is a “shift” representing the fraction of ARC target size (grep -w c /proc/spl/kstat/zfs/arcstats). The ARC target size is shifted to the right. Thus a value of ‘2’ results in the fraction = 1/4, while a value of ‘5’ results in the fraction = 1/32. Performance tuning of dbuf cache can be monitored using: - dbufstat command - node_exporter ZFS module for prometheus environments - telegraf ZFS plugin for general-purpose metric collection - /proc/spl/kstat/zfs/dbufstats kstat

dbuf_metadata_cache_max_bytes

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.2 - master

UINT64_MAX

v2.0 - v2.1

ULONG_MAX

v0.8

0

Range:

0 = use dbuf_metadata_cache_shift to ARC c_max

Change:

Dynamic

Tags:

dbuf, dbuf_cache, metadata

Maximum size in bytes of the metadata dbuf cache. The target size is determined by the MIN versus 1/2^dbuf_metadata_cache_shift (1/64th) of the target ARC size. The behavior of the metadata dbuf cache and its associated settings can be observed via the /proc/spl/kstat/zfs/dbufstats kstat.

Notes: dbuf_metadata_cache_max_bytes sets the size of the dbuf metadata cache as a number of bytes. This is an alternate method for setting dbuf metadata cache size than dbuf_metadata_cache_shift dbuf_metadata_cache_max_bytes overrides dbuf_metadata_cache_shift

dbuf_metadata_cache_shift

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

6

Range:

practical range is ( dbuf_cache_shift + 1) to MAX_INT

Change:

Dynamic

Tags:

ARC, dbuf, dbuf_cache, metadata

Set the size of the dbuf metadata cache (dbuf_metadata_cache_max_bytes) to a log2 fraction of the target ARC size.

Notes: dbuf_metadata_cache_shift sets the size of the dbuf metadata cache as a fraction of ARC target size. This is an alternate method for setting dbuf metadata cache size than dbuf_metadata_cache_max_bytes. dbuf_metadata_cache_max_bytes overrides dbuf_metadata_cache_shift This value is a “shift” representing the fraction of ARC target size (grep -w c /proc/spl/kstat/zfs/arcstats). The ARC target size is shifted to the right. Thus a value of ‘2’ results in the fraction = 1/4, while a value of ‘6’ results in the fraction = 1/64.

dbuf_mutex_cache_shift

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Change:

Prior to module load

Tags:

dbuf

Set the size of the mutex array for the dbuf cache. When set to 0 the array is dynamically sized based on total system memory.

ddt_zap_default_bs

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

15

Change:

Dynamic

Tags:

DDT, ddt_zap, dedup

Default DDT ZAP data block size as a power of 2. Note that changing this after creating a DDT on the pool will not affect existing DDTs, only newly created ones.

ddt_zap_default_ibs

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

15

Change:

Dynamic

Tags:

DDT, ddt_zap, dedup

Default DDT ZAP indirect block size as a power of 2. Note that changing this after creating a DDT on the pool will not affect existing DDTs, only newly created ones.

dmu_ddt_copies

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

3

Change:

Dynamic

Tags:

DMU

Controls the number of copies stored for DeDup Table (DDT) objects. Reducing the number of copies to 1 from the previous default of 3 can reduce the write inflation caused by deduplication. This assumes redundancy for this data is provided by the vdev layer. If the DDT is damaged, space may be leaked (not freed) when the DDT can not report the correct reference count.

dmu_object_alloc_chunk_shift

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

7

Range:

7 to 9

Change:

Dynamic

Tags:

allocation, DMU

dnode slots allocated in a single operation as a power of 2. The default value minimizes lock contention for the bulk operation performed.

When to change: If the workload creates many files concurrently on a system with many CPUs, then increasing dmu_object_alloc_chunk_shift can decrease lock contention

Notes: Each of the concurrent object allocators grabs 2^dmu_object_alloc_chunk_shift dnode slots at a time. The default is to grab 128 slots, or 4 blocks worth. This default value was experimentally determined to be the lowest value that eliminates the measurable effect of lock contention in the DMU object allocation code path.

dmu_prefetch_max

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

134217728

Change:

Dynamic

Tags:

DMU, prefetch

Limit the amount we can prefetch with one call to this amount in bytes. This helps to limit the amount of memory that can be used by prefetching.

icp_aes_impl

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Range:

varies by hardware

Change:

Dynamic

Tags:

encryption, ICP

Select aes implementation.

When to change: debugging ZFS encryption on hardware

Notes: By default, ZFS will choose the highest performance, hardware-optimized implementation of the AES encryption algorithm. The icp_aes_impl tunable overrides this automatic choice. Note: icp_aes_impl is set in the icp kernel module, not the zfs kernel module. To observe the available options cat /sys/module/icp/parameters/icp_aes_impl The default option is shown in brackets ‘[]’

icp_gcm_avx_chunk_size

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Change:

Dynamic

Tags:

gcm, ICP

How many bytes to process while owning the FPU

icp_gcm_impl

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Range:

varies by hardware

Change:

Dynamic

Tags:

encryption, gcm, ICP

Select gcm implementation.

When to change: debugging ZFS encryption on hardware

Notes: By default, ZFS will choose the highest performance, hardware-optimized implementation of the GCM encryption algorithm. The icp_gcm_impl tunable overrides this automatic choice. Note: icp_gcm_impl is set in the icp kernel module, not the zfs kernel module. To observe the available options cat /sys/module/icp/parameters/icp_gcm_impl The default option is shown in brackets ‘[]’

ignore_hole_birth

Versions:

v0.7 - v2.2

Platforms:

Linux, FreeBSD

Type:

int

Range:

0

disabled

1

enabled

Change:

Dynamic

Tags:

DMU, dmu_traverse, send

Alias for send_holes_without_birth_time.

When to change: Enable if you suspect your datasets are affected by a bug in hole_birth during zfs send operations

Notes: Renamed to send_holes_without_birth_time in v2.4.0 (commit 284580c87). When set, the hole_birth optimization will not be used and all holes will always be sent by zfs send.

l2arc_dwpd_limit

Versions:

master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

100

Change:

Dynamic

Tags:

ARC, L2ARC

Drive Writes Per Day limit for L2ARC devices to protect SSD endurance, specified as a percentage where 100 equals 1.0 DWPD. A value of 100 means each L2ARC device can write its own capacity once per day. Lower values support fractional DWPD (50 = 0.5 DWPD, 30 = 0.3 DWPD for QLC SSDs). Higher values allow more writes (300 = 3.0 DWPD). The effective write rate is always bounded by l2arc_write_max. A value of 0 disables DWPD rate limiting entirely. DWPD limiting only applies after the initial fill pass completes and when total L2ARC capacity is at least twice arc_c_max.

l2arc_exclude_special

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

disabled

1

enabled

Change:

Dynamic

Tags:

ARC, L2ARC, special_vdev

Controls whether buffers present on special vdevs are eligible for caching into L2ARC. If set to 1, exclude dbufs on special vdevs from being cached to L2ARC.

When to change: If cache and special devices exist and caching data on special devices in L2ARC is not desired

l2arc_ext_headroom_pct

Versions:

master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

25

Change:

Dynamic

Tags:

ARC, L2ARC

Percentage of each ARC state’s size that a pass may scan before resetting its markers to the tail. Lower values keep the marker closer to the tail under active workloads. Set to 0 to disable the depth cap.

l2arc_feed_again

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

disabled

1

enabled

Change:

Dynamic

Tags:

ARC, L2ARC

Turbo L2ARC warm-up. When the L2ARC is cold the fill interval will be set as fast as possible.

When to change: If cache devices exist and it is desired to fill them as fast as possible

Notes: Turbo L2ARC cache warm-up. When the L2ARC is cold the fill interval will be set to aggressively fill as fast as possible.

l2arc_feed_min_ms

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

200

Range:

0 to (1000 * l2arc_feed_secs)

Change:

Dynamic

Tags:

ARC, L2ARC

Min feed interval in milliseconds. Requires l2arc_feed_again=1 and only applicable in related situations.

When to change: If cache devices exist and l2arc_feed_again and the feed is too aggressive, then this tunable can be adjusted to reduce the impact of the fill

Notes: Minimum time period for aggressively feeding the L2ARC. The L2ARC feed thread wakes up once per second (see l2arc_feed_secs) to look for data to feed into the L2ARC. l2arc_feed_min_ms only affects the turbo L2ARC cache warm-up and allows the aggressiveness to be adjusted.

l2arc_feed_secs

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1

Range:

1 to UINT64_MAX

Change:

Dynamic

Tags:

ARC, L2ARC

Seconds between L2ARC writing.

When to change: Do not change

Notes: Seconds between waking the L2ARC feed thread. One feed thread works for all cache devices in turn. If the pool that owns a cache device is imported readonly, then the feed thread is delayed 5 * l2arc_feed_secs before moving onto the next cache device. If multiple pools are imported with cache devices and one pool with cache is imported readonly, the L2ARC feed rate to all caches can be slowed.

l2arc_headroom

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.2 - master

8

v0.6 - v2.1

2

Change:

Dynamic

Tags:

ARC, L2ARC

How far through the ARC lists to search for L2ARC cacheable content per cycle, expressed as a multiplier of the effective write size. Setting to 0 disables the per-cycle headroom limit. Scan depth is also bounded by l2arc_ext_headroom_pct when persistent markers are active.

When to change: If the rate of change in the ARC is faster than the overall L2ARC feed rate, then increasing l2arc_headroom can increase L2ARC efficiency. Setting the value too large can cause the L2ARC feed thread to consume more CPU time looking for data to feed. Setting to 0 disables the headroom limit, allowing the full length of ARC lists to be searched for cacheable content.

Notes: How far through the ARC lists to search for L2ARC cacheable content, expressed as a multiplier of l2arc_write_max

l2arc_headroom_boost

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

200

Range:

100 to UINT64_MAX, when set to 100, the L2ARC headroom boost feature is effectively disabled

Change:

Dynamic

Tags:

ARC, L2ARC

Scales l2arc_headroom by this percentage when L2ARC contents are being successfully compressed before writing. A value of 100 disables this feature.

When to change: If average compression efficiency is greater than 2:1, then increasing l2arc_headroom_boost can increase the L2ARC feed rate

Notes: Percentage scale for l2arc_headroom when L2ARC contents are being successfully compressed before writing.

l2arc_meta_cycles

Versions:

master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

2

Change:

Dynamic

Tags:

ARC, L2ARC

How many consecutive cycles metadata may monopolize the write budget before being skipped to let data run. The default of 2 gives metadata roughly 67% and data 33% of L2ARC write bandwidth under sustained load. Higher values favor metadata; set to 0 to disable.

l2arc_meta_percent

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

33

Range:

0 to 100

Change:

Dynamic

Tags:

ARC, L2ARC

Percent of ARC size allowed for L2ARC-only headers. Since L2ARC buffers are not evicted on memory pressure, too many headers on a system with an irrationally large L2ARC can render it slow or unusable. This parameter limits L2ARC writes and rebuilds to achieve the target.

When to change: When workload really require enormous L2ARC.

Notes: Percent of ARC size allowed for L2ARC-only headers. Since L2ARC buffers are not evicted on memory pressure, too large amount of headers on system with irrationally large L2ARC can render it slow or unusable. This parameter limits L2ARC writes and rebuild to achieve it.

l2arc_mfuonly

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1 | 2

Change:

Dynamic

Tags:

ARC, L2ARC

Controls whether only MFU metadata and data are cached from ARC into L2ARC. This may be desired to avoid wasting space on L2ARC when reading/writing large amounts of data that are not expected to be accessed more than once.

The default is 0, meaning both MRU and MFU data and metadata are cached. When turning off this feature (setting it to 0), some MRU buffers will still be present in ARC and eventually cached on L2ARC. If l2arc_noprefetch=0, some prefetched buffers will be cached to L2ARC, and those might later transition to MRU, in which case the l2arc_mru_asize arcstat will not be 0.

Setting it to 1 means to L2 cache only MFU data and metadata.

Setting it to 2 means to L2 cache all metadata (MRU+MFU) but only MFU data (i.e. MRU data are not cached). This can be the right setting to cache as much metadata as possible even when having high data turnover.

Regardless of l2arc_noprefetch, some MFU buffers might be evicted from ARC, accessed later on as prefetches and transition to MRU as prefetches. If accessed again they are counted as MRU and the l2arc_mru_asize arcstat will not be 0.

The ARC status of L2ARC buffers when they were first cached in L2ARC can be seen in the l2arc_mru_asize, l2arc_mfu_asize, and l2arc_prefetch_asize arcstats when importing the pool or onlining a cache device if persistent L2ARC is enabled.

The evict_l2_eligible_mru arcstat does not take into account if this option is enabled as the information provided by the evict_l2_eligible_m[rf]u arcstats can be used to decide if toggling this option is appropriate for the current workload.

When to change: When accessing a large amount of data only once.

l2arc_nocompress

Versions:

v0.6

Platforms:

Linux, FreeBSD

Type:

int

Range:

0

store compressed blocks in cache device

1

store uncompressed blocks in cache device

Change:

Dynamic

Tags:

ARC, L2ARC

Skip compressing L2ARC buffers

Use 1 for yes and 0 for no (default).

When to change: When testing compressed L2ARC feature

Notes: Removed in v0.7.0 by commit d3c2ae1c0 (“OpenZFS 6950 - ARC should cache compressed data”). The ARC now always caches compressed data, making this toggle unnecessary. No replacement parameter. Disable writing compressed data to cache devices. Disabling allows the legacy behavior of writing decompressed data to cache devices.

l2arc_noprefetch

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

write prefetched but unused buffers to cache devices

1

do not write prefetched but unused buffers to cache devices

Change:

Dynamic

Tags:

ARC, L2ARC, prefetch

Do not write buffers to L2ARC if they were prefetched but not used by applications. In case there are prefetched buffers in L2ARC and this option is later set, we do not read the prefetched buffers from L2ARC. Unsetting this option is useful for caching sequential reads from the disks to L2ARC and serve those reads from L2ARC later on. This may be beneficial in case the L2ARC device is significantly faster in sequential reads than the disks of the pool.

Use 1 to disable and 0 to enable caching/reading prefetches to/from L2ARC.

When to change: Setting to 0 can increase L2ARC hit rates for workloads where the ARC is too small for a read workload that benefits from prefetching. Also, if the main pool devices are very slow, setting to 0 can improve some workloads such as backups.

l2arc_norw

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

read and write simultaneously

1

avoid writes when reading for antique SSDs

Change:

Dynamic

Tags:

ARC, L2ARC

No reads during writes.

When to change: In the early days of SSDs, some devices did not perform well when reading and writing simultaneously. Modern SSDs do not have these issues.

Notes: Disables writing to cache devices while they are being read.

l2arc_rebuild_blocks_min_l2size

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1073741824

Change:

Dynamic

Tags:

ARC, L2ARC

Minimum size of an L2ARC device required in order to write log blocks in it. The log blocks are used upon importing the pool to rebuild the persistent L2ARC.

For L2ARC devices less than 1 GiB, the amount of data l2arc_evict() evicts is significant compared to the amount of restored L2ARC data. In this case, do not write log blocks in L2ARC in order not to waste space.

When to change: The cache device is small and the pool is frequently imported.

Notes: The minimum required size (in bytes) of an L2ARC device in order to write log blocks in it. The log blocks are used upon importing the pool to rebuild the persistent L2ARC. For L2ARC devices less than 1GB the overhead involved offsets most of benefit so log blocks are not written for cache devices smaller than this.

l2arc_rebuild_enabled

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

disable persistent L2ARC rebuild

1

enable persistent L2ARC rebuild

Change:

Dynamic

Tags:

ARC, L2ARC

Rebuild the L2ARC when importing a pool (persistent L2ARC). This can be disabled if there are problems importing a pool or attaching an L2ARC device (e.g. the L2ARC device is slow in reading stored log metadata, or the metadata has become somehow fragmented/unusable).

When to change: If there are problems importing a pool or attaching an L2ARC device.

l2arc_trim_ahead

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

0

Range:

0 to 100

Change:

Dynamic

Tags:

ARC, L2ARC, trim

Trims ahead of the current write size on L2ARC devices by this percentage of write size if we have filled the device. If set to 100 we TRIM twice the space required to accommodate upcoming writes. A minimum of 64 MiB will be trimmed. It also enables TRIM of the whole L2ARC device upon creation or addition to an existing pool or if the header of the device is invalid upon importing a pool or onlining a cache device. A value of 0 disables TRIM on L2ARC altogether and is the default as it can put significant stress on the underlying storage devices. This will vary depending of how well the specific device handles these commands.

When to change: Consider setting for cache devices which efficiently handle TRIM commands.

Notes: Once the cache device has been filled TRIM ahead of the current write size l2arc_write_max on L2ARC devices by this percentage. This can speed up future writes depending on the performance characteristics of the cache device. When set to 100% TRIM twice the space required to accommodate upcoming writes. A minimum of 64MB will be trimmed. If set it enables TRIM of the whole L2ARC device when it is added to a pool. By default, this option is disabled since it can put significant stress on the underlying storage devices.

l2arc_write_boost

Versions:

v0.6 - v2.4

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.2 - v2.4

33554432

v0.6 - v2.1

8,388,608

Change:

Dynamic

Tags:

ARC, L2ARC

Cold L2ARC devices will have l2arc_write_max increased by this amount while they remain cold.

When to change: To fill the cache devices more aggressively after pool import.

Notes: Until the ARC fills, increases the L2ARC fill rate l2arc_write_max by l2arc_write_boost.

l2arc_write_max

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.2 - master

33554432

v0.6 - v2.1

8,388,608

Range:

1 to UINT64_MAX

Change:

Dynamic

Tags:

ARC, L2ARC

Maximum write rate in bytes per second for each L2ARC device. Used directly during initial fill, when DWPD limiting is disabled, or for non-persistent L2ARC. When DWPD limiting is active, writes are capped by this rate. Total L2ARC throughput scales with the number of cache devices in a pool.

When to change: If the cache devices can sustain the write workload, increasing the rate of cache device fill when workloads generate new data at a rate higher than l2arc_write_max can increase L2ARC hit rate

Notes: Maximum number of bytes to be written to each cache device for each L2ARC feed thread interval (see l2arc_feed_secs). Until v2.4 the limit could be raised during warmup by l2arc_write_boost, which has since been removed. By default l2arc_feed_secs is 1 second, delivering a maximum write workload to cache devices of 32 MiB/sec.

metaslab_aliquot

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.4 - master

2097152

v2.1 - v2.3

1048576

v0.6 - v2.0

524,288

Change:

Dynamic

Tags:

allocation, metaslab, vdev

Metaslab group’s per child vdev allocation granularity, in bytes. This is roughly similar to what would be referred to as the “stripe size” in traditional RAID arrays. In normal operation, ZFS will try to write this amount of data to each child of a top-level vdev before moving on to the next top-level vdev.

When to change: If write performance increases as devices more efficiently write larger, contiguous blocks

Notes: Sets the metaslab granularity. Nominally, ZFS will try to allocate this amount of data to a top-level vdev before moving on to the next top-level vdev. This is roughly similar to what would be referred to as the “stripe size” in traditional RAID arrays. When tuning for HDDs, it can be more efficient to have a few larger, sequential writes to a device rather than switching to the next device. Monitoring the size of contiguous writes to the disks relative to the write throughput can be used to determine if increasing metaslab_aliquot can help. For modern devices, it is unlikely that decreasing metaslab_aliquot from the default will help. If there is only one top-level vdev, this tunable is not used.

metaslab_bias_enabled

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

spread evenly across top-level vdevs

1

bias spread to favor less full top-level vdevs

Change:

Dynamic

Tags:

allocation, metaslab, vdev

Enable metaslab groups biasing based on their over- or under-utilization relative to the metaslab class average. If disabled, each metaslab group will receive allocations proportional to its capacity.

When to change: If a new top-level vdev is added and you do not want to bias new allocations to the new top-level vdev

Notes: Enables metaslab group biasing based on a top-level vdev’s utilization relative to the pool. Nominally, all top-level devs are the same size and the allocation is spread evenly. When the top-level vdevs are not of the same size, for example if a new (empty) top-level is added to the pool, this allows the new top-level vdev to get a larger portion of new allocations.

metaslab_debug_load

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

do not load all metaslab info at pool import

1

dynamically load metaslab info as needed

Change:

Dynamic

Tags:

allocation, debug, memory, metaslab

Load all metaslabs during pool import.

When to change: When RAM is plentiful and pool import time is not a consideration

Notes: When enabled, all metaslabs are loaded into memory during pool import. Nominally, metaslab space map information is loaded and unloaded as needed (see metaslab_debug_unload) It is difficult to predict how much RAM is required to store a space map. An empty or completely full metaslab has a small space map. However, a highly fragmented space map can consume significantly more memory. Enabling metaslab_debug_load can increase pool import time.

metaslab_debug_unload

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

dynamically unload metaslab info

1

unload metaslab info only upon pool export

Change:

Dynamic

Tags:

allocation, debug, memory, metaslab

Prevent metaslabs from being unloaded.

When to change: When RAM is plentiful and the penalty for dynamically reloading metaslab info from the pool is high

Notes: When enabled, prevents metaslab information from being dynamically unloaded from RAM. Nominally, metaslab space map information is loaded and unloaded as needed (see metaslab_debug_load) It is difficult to predict how much RAM is required to store a space map. An empty or completely full metaslab has a small space map. However, a highly fragmented space map can consume significantly more memory. Enabling metaslab_debug_unload consumes RAM that would otherwise be freed.

metaslab_df_alloc_threshold

Versions:

master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

131072

Change:

Dynamic

Tags:

metaslab

Minimum size which forces the dynamic allocator to change its allocation strategy. Once the space map cannot satisfy an allocation of this size, it switches to a more aggressive strategy (searching by size rather than offset).

metaslab_df_free_pct

Versions:

master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

4

Change:

Dynamic

Tags:

metaslab

The minimum free space, in percent, which must be available in a space map to continue allocations in a first-fit fashion. Once free space drops below this level, allocations switch to a best-fit strategy.

metaslab_df_use_largest_segment

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

metaslab

If not searching forward (due to metaslab_df_max_search, metaslab_df_free_pct, or metaslab_df_alloc_threshold), this tunable controls which segment is used. If set, we will use the largest free segment. If unset, we will use a segment of at least the requested size.

metaslab_force_ganging

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

16777217

Range:

SPA_MINBLOCKSIZE to (SPA_MAXBLOCKSIZE + 1)

Change:

Dynamic

Tags:

allocation, metaslab

Make some blocks above a certain size be gang blocks. This option is used by the test suite to facilitate testing.

When to change: for development testing purposes only

Notes: When testing allocation code, metaslab_force_ganging forces blocks above the specified size to be ganged.

metaslab_force_ganging_pct

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

3

Change:

Dynamic

Tags:

metaslab

For blocks that could be forced to be a gang block (due to metaslab_force_ganging), force this many of them to be gang blocks.

metaslab_fragmentation_factor_enabled

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not consider metaslab free space fragmentation

1

try to avoid fragmented metaslabs

Change:

Dynamic

Tags:

allocation, metaslab

Enable use of the fragmentation metric in computing metaslab weights.

When to change: To test metaslab fragmentation

Notes: Enable use of the fragmentation metric in computing metaslab weights. In version v0.7.0, if zfs_metaslab_segment_weight_enabled is enabled, then metaslab_fragmentation_factor_enabled is ignored.

metaslab_lba_weighting_enabled

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not use LBA weighting

1

use LBA weighting

Change:

Dynamic

Tags:

allocation, HDD, metaslab, SSD

Give more weight to metaslabs with lower LBAs, assuming they have greater bandwidth, as is typically the case on a modern constant angular velocity disk drive.

When to change: disable if using only SSDs and version v0.6.4 or earlier

Verification: The rotational setting described by a block device in sysfs by observing /sys/ block/DISK_NAME/queue/rotational

Notes: Modern HDDs have uniform bit density and constant angular velocity. Therefore, the outer recording zones are faster (higher bandwidth) than the inner zones by the ratio of outer to inner track diameter. The difference in bandwidth can be 2:1, and is often available in the HDD detailed specifications or drive manual. For HDDs when metaslab_lba_weighting_enabled is true, write allocation preference is given to the metaslabs representing the outer recording zones. Thus the allocation to metaslabs prefers faster bandwidth over free space. If the devices are not rotational, yet misrepresent themselves to the OS as rotational, then disabling metaslab_lba_weighting_enabled can result in more even, free-space-based allocation.

metaslab_perf_bias

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0 | 2

Change:

Dynamic

Tags:

metaslab

Controls metaslab groups biasing based on their write performance. Setting to 0 makes all metaslab groups receive fixed amounts of allocations. Setting to 2 allows faster metaslab groups to allocate more. Setting to 1 equals to 2 if the pool is write-bound or 0 otherwise. That is, if the pool is limited by write throughput, then allocate more from faster metaslab groups, but if not, try to evenly distribute the allocations.

metaslab_preload_enabled

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not preload metaslab info

1

preload up to 3 metaslabs

Change:

Dynamic

Tags:

allocation, metaslab

Enable metaslab group preloading.

When to change: When testing metaslab allocation

Notes: Enable metaslab group preloading. Each top-level vdev has a metaslab group. By default, up to 3 copies of metadata can exist and are distributed across multiple top-level vdevs. metaslab_preload_enabled allows the corresponding metaslabs to be preloaded, thus improving allocation efficiency.

metaslab_preload_limit

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Change:

Dynamic

Tags:

metaslab

Maximum number of metaslabs per group to preload

metaslab_preload_pct

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

50

Change:

Dynamic

Tags:

metaslab, SPA

Percentage of CPUs to run a metaslab preload taskq

metaslab_trace_enabled

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Change:

Dynamic

Tags:

metaslab

Enable metaslab allocation tracing

metaslab_trace_max_entries

Versions:

master

Platforms:

Linux, FreeBSD

Type:

u64

Change:

Dynamic

Tags:

metaslab

Maximum entries for metaslab allocation tracing

metaslab_unload_delay

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

32

Change:

Dynamic

Tags:

delay, metaslab

After a metaslab is used, we keep it loaded for this many TXGs, to attempt to reduce unnecessary reloading. Note that both this many TXGs and metaslab_unload_delay_ms milliseconds must pass before unloading will occur.

metaslab_unload_delay_ms

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

600000

Change:

Dynamic

Tags:

delay, metaslab

After a metaslab is used, we keep it loaded for this many milliseconds, to attempt to reduce unnecessary reloading. Note, that both this many milliseconds and metaslab_unload_delay TXGs must pass before unloading will occur.

metaslabs_per_vdev

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Default:

200

Range:

16 to UINT64_MAX

Change:

Dynamic

Tags:

allocation, metaslab, vdev

When a vdev is added, it will be divided into approximately (but no more than) this number of metaslabs.

Default value: 200.

When to change: When testing metaslab allocation

raidz_expand_max_copy_bytes

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

160MB

Change:

Dynamic

Tags:

raidz, vdev

Max amount of memory to use for RAID-Z expansion I/O. This limits how much I/O can be outstanding at once.

raidz_expand_max_reflow_bytes

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

0

Change:

Dynamic

Tags:

raidz, vdev

For testing, pause RAID-Z expansion when reflow amount reaches this value.

raidz_io_aggregate_rows

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

4

Change:

Dynamic

Tags:

raidz, vdev

For expanded RAID-Z, aggregate reads that have more rows than this.

reference_history

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

3

Change:

Dynamic

Tags:

refcount

Maximum reference holders being tracked when reference_tracking_enable is active.

reference_tracking_enable

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

refcount

Track reference holders to refcount_t objects (debug builds only).

send_holes_without_birth_time

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

DMU, dmu_traverse, send

When set, the hole_birth optimization will not be used, and all holes will always be sent during a zfs send. This is useful if you suspect your datasets are affected by a bug in hole_birth.

Notes: Renamed from ignore_hole_birth in v2.4.0.

spa_asize_inflation

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

24

Range:

1 to 24

Change:

Dynamic

Tags:

allocation, SPA

Multiplication factor used to estimate actual disk consumption from the size of data being written. The default value is a worst case estimate, but lower values may be valid for a given pool depending on its configuration. Pool administrators who understand the factors involved may wish to specify a more realistic inflation factor, particularly if they operate close to quota or capacity limits.

When to change: If the allocation requirements for the workload are well known and quotas are used

Notes: Multiplication factor used to estimate actual disk consumption from the size of data being written. The default value is a worst case estimate, but lower values may be valid for a given pool depending on its configuration. Pool administrators who understand the factors involved may wish to specify a more realistic inflation factor, particularly if they operate close to quota or capacity limits. The worst case space requirement for allocation is single-sector max-parity RAIDZ blocks, in which case the space requirement is exactly 4 times the size, accounting for a maximum of 3 parity blocks. This is added to the maximum number of ZFS copies parameter (copies max=3). Additional space is required if the block could impact deduplication tables. Altogether, the worst case is 24. If the estimation is not correct, then quotas or out-of-space conditions can lead to optimistic expectations of the ability to allocate. Applications are typically not prepared to deal with such failures and can misbehave.

spa_config_path

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

string

Change:

Prior to module load

Tags:

import, SPA, spa_config

SPA config file.

When to change: If creating a non-standard distribution and the cachefile property is inconvenient

Notes: By default, the zpool import command searches for pool information in the zpool.cache file. If the pool to be imported has an entry in zpool.cache then the devices do not have to be scanned to determine if they are pool members. The path to the cache file is spa_config_path. For more information on zpool import and the -o cachefile and -d options, see the man page for zpool(8) See also zfs_autoimport_disable (removed after v2.3)

spa_cpus_per_allocator

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

4

Change:

Dynamic

Tags:

SPA

Determines the minimum number of CPUs in a system for block allocator per spa instance. Set value only applies to pools imported/created after that.

spa_flush_txg_time

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

uint

Change:

Dynamic

Tags:

SPA

How frequently the TXG timestamps database should be flushed to disk (in seconds)

spa_load_print_vdev_tree

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

do not print pool configuration in logs

1

print pool configuration in logs

Change:

Dynamic

Tags:

import, SPA, vdev

Whether to print the vdev tree in the debugging message buffer during pool import.

When to change: troubleshooting pool import failures

Notes: spa_load_print_vdev_tree enables printing of the attempted pool import’s vdev tree to kernel message to the ZFS debug message log /proc/spl/kstat/zfs/dbgmsg Both the provided vdev tree and MOS vdev tree are printed, which can be useful for debugging problems with the zpool cachefile

spa_load_verify_data

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not verify data upon pool import

1

verify pool data upon import

Change:

Dynamic

Tags:

allocation, SPA

Whether to traverse data blocks during an “extreme rewind” (-X) import.

Data blocks are traversed only if the number of errors found in them is going to be reported, like during a dry run (-nX), since an actual rewind tolerates any number of them. If this parameter is unset, the traversal skips non-metadata blocks. It can be toggled once the import has started to stop or start the traversal of non-metadata blocks.

When to change: At the risk of data integrity, to speed extreme import of large pool

Notes: An extreme rewind import (see zpool import -X) normally performs a full traversal of all blocks in the pool for verification. If this parameter is set to 0, the traversal skips non-metadata blocks. It can be toggled once the import has started to stop or start the traversal of non-metadata blocks. See also spa_load_verify_metadata.

spa_load_verify_maxinflight

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Default:

10000

Range:

1 to MAX_INT

Change:

Dynamic

Tags:

import, SPA

Maximum concurrent I/Os during the traversal performed during an “extreme rewind” (-X) pool import.

Default value: 10000.

When to change: During an extreme rewind import, to match the concurrent I/O capabilities of the pool devices

Notes: Maximum number of concurrent I/Os during the data verification performed during an extreme rewind import (see zpool import -X)

spa_load_verify_metadata

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not verify metadata upon pool import

1

verify pool metadata upon import

Change:

Dynamic

Tags:

import, metadata, SPA

Whether to traverse blocks during an “extreme rewind” (-X) pool import.

An extreme rewind import normally performs a full traversal of all blocks in the pool for verification. If this parameter is unset, the traversal is not performed. It can be toggled once the import has started to stop or start the traversal.

When to change: At the risk of data integrity, to speed extreme import of large pool

Notes: An extreme rewind import (see zpool import -X) normally performs a full traversal of all blocks in the pool for verification. If this parameter is set to 0, the traversal is not performed. It can be toggled once the import has started to stop or start the traversal. See spa_load_verify_data

spa_load_verify_shift

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

4

Range:

1 to MAX_INT

Change:

Dynamic

Tags:

ARC, import, SPA

Sets the maximum number of bytes to consume during pool import to the log2 fraction of the target ARC size.

When to change: troubleshooting pool import on large memory machines

Notes: spa_load_verify_shift sets the fraction of ARC that can be used by inflight I/Os when verifying the pool during import. This value is a “shift” representing the fraction of ARC target size (grep -w c /proc/spl/kstat/zfs/arcstats). The ARC target size is shifted to the right. Thus a value of ‘2’ results in the fraction = 1/4, while a value of ‘4’ results in the fraction = 1/8. For large memory machines, pool import can consume large amounts of ARC: much larger than the value of maxinflight. This can result in spa_load_verify_maxinflight (removed after v0.7) having a value of 0 causing the system to hang. Setting spa_load_verify_shift can reduce this limit and allow importing without hanging.

spa_note_txg_time

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

uint

Change:

Dynamic

Tags:

SPA

How frequently TXG timestamps are stored internally (in seconds)

spa_num_allocators

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

4

Change:

Dynamic

Tags:

SPA

Determines the number of block allocators to use per spa instance. Capped by the number of actual CPUs in the system via spa_cpus_per_allocator.

Note that setting this value too high could result in performance degradation and/or excess fragmentation. Set value only applies to pools imported/created after that.

spa_slop_shift

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

5

Range:

1 to MAX_INT, however the practical upper limit is 15 for a system with 4TB of RAM

Change:

Dynamic

Tags:

allocation, SPA

Normally, we don’t allow the last 3.2% (1/2^spa_slop_shift) of space in the pool to be consumed. This ensures that we don’t run the pool completely out of space, due to unaccounted changes (e.g. to the MOS). It also limits the worst-case time to allocate space. If we have less than this amount of free space, most ZPL operations (e.g. write, create) will return ENOSPC.

When to change: For large pools, when 3.2% may be too conservative and more usable space is desired, consider increasing spa_slop_shift

Notes: Normally, the last 3.2% (1/(2^spa_slop_shift)) of pool space is reserved to ensure the pool doesn’t run completely out of space, due to unaccounted changes (e.g. to the MOS). This also limits the worst-case time to allocate space. When less than this amount of free space exists, most ZPL operations (e.g. write, create) return error:no space (ENOSPC). Changing spa_slop_shift affects the currently loaded ZFS module and all imported pools. spa_slop_shift is not stored on disk. Beware when importing full pools on systems with larger spa_slop_shift can lead to over-full conditions. The minimum SPA slop space is limited to 128 MiB. The maximum SPA slop space is limited to 128 GiB.

spa_upgrade_errlog_limit

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Change:

Dynamic

Tags:

SPA

Limits the number of on-disk error log entries that will be converted to the new format when enabling the head_errlog feature. The default is to convert all log entries.

spl_hostid

Versions:

v0.8 - master

Platforms:

Linux

Type:

ulong

Default:

0

Range:

0=ignore hostid, 1 to 4,294,967,295 (32-bits or 0xffffffff)

Change:

Dynamic

Tags:

generic, hostid, MMP, SPL

The system hostid, when set this can be used to uniquely identify a system. By default this value is set to zero which indicates the hostid is disabled. It can be explicitly enabled by placing a unique non-zero value in /etc/hostid.

When to change: to uniquely identify a system when vdevs can be shared across multiple systems

Notes: spl_hostid is a unique system id number. It originated in Sun’s products where most systems had a unique id assigned at the factory. This assignment does not exist in modern hardware. In ZFS, the hostid is stored in the vdev label and can be used to determine if another system had imported the pool. When set spl_hostid can be used to uniquely identify a system. By default this value is set to zero which indicates the hostid is disabled. It can be explicitly enabled by placing a unique non-zero value in the file shown in spl_hostid_path

spl_hostid_path

Versions:

v0.8 - master

Platforms:

Linux

Type:

charp

Change:

Prior to module load

Tags:

generic, hostid, MMP, SPL

The expected path to locate the system hostid when specified. This value may be overridden for non-standard configurations.

When to change: when creating a new ZFS distribution where the default value is inappropriate

Notes: spl_hostid_path is the path name for a file that can contain a unique hostid. For testing purposes, spl_hostid_path can be overridden by the ZFS_HOSTID environment variable.

spl_kmem_alloc_max

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

v2.1 - master

KMALLOC_MAX_SIZE

v0.8 - v2.0

KMALLOC_MAX_SIZE/4

Change:

Dynamic

Tags:

kmem, memory, SPL

Large kmem_alloc() allocations will fail if they exceed KMALLOC_MAX_SIZE. Allocations which are marginally smaller than this limit may succeed but should still be avoided due to the expense of locating a contiguous range of free pages. Therefore, a maximum kmem size with reasonable safely margin of 4x is set. kmem_alloc() allocations larger than this maximum will quickly fail. vmem_alloc() allocations less than or equal to this value will use kmalloc(), but shift to vmalloc() when exceeding this value.

Notes: Large kmem_alloc() allocations fail if they exceed KMALLOC_MAX_SIZE, as determined by the kernel source. Allocations which are marginally smaller than this limit may succeed but should still be avoided due to the expense of locating a contiguous range of free pages. Therefore, a maximum kmem size with reasonable safely margin of 4x is set. kmem_alloc() allocations larger than this maximum will quickly fail. vmem_alloc() allocations less than or equal to this value will use kmalloc(), but shift to vmalloc() when exceeding this value.

spl_kmem_alloc_warn

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

32768

Range:

0=disable the warnings,

Change:

Dynamic

Tags:

kmem, memory, SPL

As a general rule kmem_alloc() allocations should be small, preferably just a few pages, since they must by physically contiguous. Therefore, a rate limited warning will be printed to the console for any kmem_alloc() which exceeds a reasonable threshold.

The default warning threshold is set to eight pages but capped at 32K to accommodate systems using large pages. This value was selected to be small enough to ensure the largest allocations are quickly noticed and fixed. But large enough to avoid logging any warnings when a allocation size is larger than optimal but not a serious concern. Since this value is tunable, developers are encouraged to set it lower when testing so any new largish allocations are quickly caught. These warnings may be disabled by setting the threshold to zero.

When to change: developers are encouraged lower when testing so any new, large allocations are quickly caught

Notes: As a general rule kmem_alloc() allocations should be small, preferably just a few pages since they must by physically contiguous. Therefore, a rate limited warning is printed to the console for any kmem_alloc() which exceeds the threshold spl_kmem_alloc_warn The default warning threshold is set to eight pages but capped at 32K to accommodate systems using large pages. This value was selected to be small enough to ensure the largest allocations are quickly noticed and fixed. But large enough to avoid logging any warnings when a allocation size is larger than optimal but not a serious concern. Since this value is tunable, developers are encouraged to set it lower when testing so any new largish allocations are quickly caught. These warnings may be disabled by setting the threshold to zero.

spl_kmem_cache_expire

Versions:

v0.8

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0x02

Range:

0x01 - Aging (illumos), 0x02 - Low memory (Linux)

Change:

Dynamic

Tags:

kmem, kmem_cache, memory, SPL

Cache expiration is part of default Illumos cache behavior. The idea is that objects in magazines which have not been recently accessed should be returned to the slabs periodically. This is known as cache aging and when enabled objects will be typically returned after 15 seconds.

On the other hand Linux slabs are designed to never move objects back to the slabs unless there is memory pressure. This is possible because under Linux the cache will be notified when memory is low and objects can be released.

By default only the Linux method is enabled. It has been shown to improve responsiveness on low memory systems and not negatively impact the performance of systems with more memory. This policy may be changed by setting the spl_kmem_cache_expire bit mask as follows, both policies may be enabled concurrently.

0x01 - Aging (Illumos), 0x02 - Low memory (Linux)

Default value: 0x02

Notes: Cache expiration is part of default illumos cache behavior. The idea is that objects in magazines which have not been recently accessed should be returned to the slabs periodically. This is known as cache aging and when enabled objects will be typically returned after 15 seconds. On the other hand Linux slabs are designed to never move objects back to the slabs unless there is memory pressure. This is possible because under Linux the cache will be notified when memory is low and objects can be released. By default only the Linux method is enabled. It has been shown to improve responsiveness on low memory systems and not negatively impact the performance of systems with more memory. This policy may be changed by setting the spl_kmem_cache_expire bit mask as follows, both policies may be enabled concurrently.

spl_kmem_cache_kmem_limit

Versions:

v0.8

Platforms:

Linux, FreeBSD

Type:

uint

Default:

PAGE_SIZE/4

Change:

Dynamic

Tags:

kmem, kmem_cache, memory, SPL

Depending on the size of a cache object it may be backed by kmalloc()’d or vmalloc()’d memory. This is because the size of the required allocation greatly impacts the best way to allocate the memory.

When objects are small and only a small number of memory pages need to be allocated, ideally just one, then kmalloc() is very efficient. However, when allocating multiple pages with kmalloc() it gets increasingly expensive because the pages must be physically contiguous.

For this reason we shift to vmalloc() for slabs of large objects which which removes the need for contiguous pages. We cannot use vmalloc() in all cases because there is significant locking overhead involved. This function takes a single global lock over the entire virtual address range which serializes all allocations. Using slightly different allocation functions for small and large objects allows us to handle a wide range of object sizes.

The spl_kmem_cache_kmem_limit value is used to determine this cutoff size. One quarter the PAGE_SIZE is used as the default value because spl_kmem_cache_obj_per_slab defaults to 16. This means that at most we will need to allocate four contiguous pages.

Default value: PAGE_SIZE/4

Notes: Depending on the size of a memory cache object it may be backed by kmalloc() or vmalloc() memory. This is because the size of the required allocation greatly impacts the best way to allocate the memory. When objects are small and only a small number of memory pages need to be allocated, ideally just one, then kmalloc() is very efficient. However, allocating multiple pages with kmalloc() gets increasingly expensive because the pages must be physically contiguous. For this reason we shift to vmalloc() for slabs of large objects which which removes the need for contiguous pages. vmalloc() cannot be used in all cases because there is significant locking overhead involved. This function takes a single global lock over the entire virtual address range which serializes all allocations. Using slightly different allocation functions for small and large objects allows us to handle a wide range of object sizes. The spl_kmem_cache_kmem_limit value is used to determine this cutoff size. One quarter of the kernel’s compiled PAGE_SIZE is used as the default value because spl_kmem_cache_obj_per_slab defaults to 8. With these default values, at most two contiguous pages are allocated.

spl_kmem_cache_kmem_threads

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

4

Range:

1 to MAX_INT

Change:

Prior to module load

Tags:

CPU, kmem, kmem_cache, memory, SPL

The number of threads created for the spl_kmem_cache task queue. This task queue is responsible for allocating new slabs for use by the kmem caches. For the majority of systems and workloads only a small number of threads are required.

When to change: read-only

Notes: spl_kmem_cache_kmem_threads shows the current number of spl_kmem_cache threads. This task queue is responsible for allocating new slabs for use by the kmem caches. For the majority of systems and workloads only a small number of threads are required.

spl_kmem_cache_magazine_size

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

0

Range:

0=automatically scale magazine size, otherwise 2 to 256

Change:

Prior to module load

Tags:

CPU, kmem, kmem_cache, memory, SPL

Cache magazines are an optimization designed to minimize the cost of allocating memory. They do this by keeping a per-CPU cache of recently freed objects, which can then be reallocated without taking a lock. This can improve performance on highly contended caches. However, because objects in magazines will prevent otherwise empty slabs from being immediately released this may not be ideal for low memory machines.

For this reason, spl_kmem_cache_magazine_size can be used to set a maximum magazine size. When this value is set to 0 the magazine size will be automatically determined based on the object size. Otherwise magazines will be limited to 2-256 objects per magazine (i.e. per CPU). Magazines may never be entirely disabled in this implementation.

Notes: spl_kmem_cache_magazine_size shows the current . Cache magazines are an optimization designed to minimize the cost of allocating memory. They do this by keeping a per-cpu cache of recently freed objects, which can then be reallocated without taking a lock. This can improve performance on highly contended caches. However, because objects in magazines will prevent otherwise empty slabs from being immediately released this may not be ideal for low memory machines. For this reason spl_kmem_cache_magazine_size can be used to set a maximum magazine size. When this value is set to 0 the magazine size will be automatically determined based on the object size. Otherwise magazines will be limited to 2-256 objects per magazine (eg per CPU). Magazines cannot be disabled entirely in this implementation.

spl_kmem_cache_max_size

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

32

Change:

Dynamic

Tags:

kmem, kmem_cache, memory, SPL

The maximum size of a kmem cache slab in MiB. This effectively limits the maximum cache object size to spl_kmem_cache_max_size/spl_kmem_cache_obj_per_slab.

Caches may not be created with object sized larger than this limit.

Notes: spl_kmem_cache_max_size is the maximum size of a kmem cache slab in MiB. This effectively limits the maximum cache object size to spl_kmem_cache_max_size / spl_kmem_cache_obj_per_slab Kmem caches may not be created with object sized larger than this limit.

spl_kmem_cache_obj_per_slab

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

8

Change:

Dynamic

Tags:

kmem, kmem_cache, memory, SPL

The preferred number of objects per slab in the cache. In general, a larger value will increase the caches memory footprint while decreasing the time required to perform an allocation. Conversely, a smaller value will minimize the footprint and improve cache reclaim time but individual allocations may take longer.

Notes: spl_kmem_cache_obj_per_slab is the preferred number of objects per slab in the kmem cache. In general, a larger value will increase the caches memory footprint while decreasing the time required to perform an allocation. Conversely, a smaller value will minimize the footprint and improve cache reclaim time but individual allocations may take longer.

spl_kmem_cache_obj_per_slab_min

Versions:

v0.8

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Change:

Dynamic

Tags:

kmem, kmem_cache, memory, SPL

The minimum number of objects allowed per slab. Normally slabs will contain spl_kmem_cache_obj_per_slab objects but for caches that contain very large objects it’s desirable to only have a few, or even just one, object per slab.

Default value: 1

When to change: debugging kmem cache operations

Notes: spl_kmem_cache_obj_per_slab_min is the minimum number of objects allowed per slab. Normally slabs will contain spl_kmem_cache_obj_per_slab objects but for caches that contain very large objects it’s desirable to only have a few, or even just one, object per slab.

spl_kmem_cache_reclaim

Versions:

v0.8 - v2.1

Platforms:

Linux

Type:

uint

Default:

0

Range:

0

enable rapid memory reclaim from kmem caches

1

disable rapid memory reclaim from kmem caches

Change:

Dynamic

Tags:

kmem, kmem_cache, memory, SPL

When this is set it prevents Linux from being able to rapidly reclaim all the memory held by the kmem caches. This may be useful in circumstances where it’s preferable that Linux reclaim memory from some other subsystem first. Setting this will increase the likelihood out of memory events on a memory constrained system.

Notes: spl_kmem_cache_reclaim prevents Linux from being able to rapidly reclaim all the memory held by the kmem caches. This may be useful in circumstances where it’s preferable that Linux reclaim memory from some other subsystem first. Setting spl_kmem_cache_reclaim increases the likelihood out of memory events on a memory constrained system.

spl_kmem_cache_slab_limit

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

16384

Change:

Dynamic

Tags:

kmem, kmem_cache, memory, SPL

For small objects the Linux slab allocator should be used to make the most efficient use of the memory. For large objects, the SPL implementation is preferred. This value is used to determine the cutoff between a small and large object. The cutoff applies to automatic cache selection. An explicit request for a Linux slab or SPL cache overrides it.

When cache selection is automatic, objects of size spl_kmem_cache_slab_limit or smaller are allocated using the Linux slab allocator, and larger objects use the SPL allocator. A cutoff of 16K was determined to be optimal for architectures using 4K pages.

Notes: For small objects the Linux slab allocator should be used to make the most efficient use of the memory. However, large objects are not supported by the Linux slab allocator and therefore the SPL implementation is preferred. spl_kmem_cache_slab_limit is used to determine the cutoff between a small and large object. Objects of spl_kmem_cache_slab_limit or smaller will be allocated using the Linux slab allocator, large objects use the SPL allocator. A cutoff of 16 KiB was determined to be optimal for architectures using 4 KiB pages.

spl_max_show_tasks

Versions:

v0.8 - v2.2

Platforms:

Linux

Type:

uint

Default:

512

Range:

0 disables the limit, 1 to MAX_UINT

Change:

Dynamic

Tags:

SPL, taskq

The maximum number of tasks per pending list in each taskq shown in /proc/spl/taskq{,-all}. Write 0 to turn off the limit. The proc file will walk the lists with lock held, reading it could cause a lock-up if the list grow too large without limiting the output. “(truncated)” will be shown if the list is larger than the limit.

Notes: spl_max_show_tasks is the limit of tasks per pending list in each taskq shown in /proc/spl/taskq and /proc/spl/taskq-all. Reading the ProcFS files walks the lists with lock held and it could cause a lock up if the list grow too large. If the list is larger than the limit, the string "(truncated)" is printed.

spl_panic_halt

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

0

Range:

0

halt thread upon assertion

1

panic kernel upon assertion

Change:

Dynamic

Tags:

debug, panic, SPL

Cause a kernel panic on assertion failures. When not enabled, the thread is halted to facilitate further debugging.

Set to a non-zero value to enable.

When to change: when debugging assertions and kernel core dumps are desired

Notes: spl_panic_halt enables kernel panic upon assertion failures. When not enabled, the asserting thread is halted to facilitate further debugging.

spl_schedule_hrtimeout_slack_us

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

0

Change:

Dynamic

Tags:

SPL

Slack value in microseconds passed to schedule_hrtimeout_range() when a condition variable times out. A non-zero value enforces the kernel coalesce the wakeup with other timers to reduce wakeup count, at the cost of some additional sleep duration. The maximum is 1000, as defined by MAX_HRTIMEOUT_SLACK_US.

Linux-only.

spl_taskq_kick

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

0

Change:

Dynamic

Tags:

SPL, taskq

Kick stuck taskq to spawn threads. When writing a non-zero value to it, it will scan all the taskqs. If any of them have a pending task more than 5 seconds old, it will kick it to spawn more threads. This can be used if you find a rare deadlock occurs because one or more taskqs didn’t spawn a thread when it should.

When to change: See description above

Notes: Upon writing a non-zero value to spl_taskq_kick, all taskqs are scanned. If any taskq has a pending task more than 5 seconds old, the taskq spawns more threads. This can be useful in rare deadlock situations caused by one or more taskqs not spawning a thread when it should.

spl_taskq_thread_bind

Versions:

v0.8 - master

Platforms:

Linux

Type:

int

Default:

0

Range:

0

taskqs are not bound to specific CPUs

1

taskqs are bound to CPUs

Change:

Dynamic

Tags:

CPU, SPL, taskq

Bind taskq threads to specific CPUs. When enabled all taskq threads will be distributed evenly across the available CPUs. By default, this behavior is disabled to allow the Linux scheduler the maximum flexibility to determine where a thread should run.

When to change: when debugging CPU scheduling options

Notes: spl_taskq_thread_bind enables binding taskq threads to specific CPUs, distributed evenly over the available CPUs. By default, this behavior is disabled to allow the Linux scheduler the maximum flexibility to determine where a thread should run.

spl_taskq_thread_dynamic

Versions:

v0.8 - master

Platforms:

Linux

Type:

int

Default:

1

Range:

0

taskq threads are not dynamic

1

taskq threads are dynamically created and destroyed

Change:

Prior to module load

Tags:

SPL, taskq

Allow dynamic taskqs. When enabled taskqs which set the TASKQ_DYNAMIC flag will by default create only a single thread. New threads will be created on demand up to a maximum allowed number to facilitate the completion of outstanding tasks. Threads which are no longer needed will be promptly destroyed. By default this behavior is enabled but it can be disabled to aid performance analysis or troubleshooting.

When to change: disable for performance analysis or troubleshooting

Notes: spl_taskq_thread_dynamic enables taskqs to set the TASKQ_DYNAMIC flag will by default create only a single thread. New threads will be created on demand up to a maximum allowed number to facilitate the completion of outstanding tasks. Threads which are no longer needed are promptly destroyed. By default this behavior is enabled but it can be d. See also zfs_zil_clean_taskq_nthr_pct, zio_taskq_batch_pct

spl_taskq_thread_priority

Versions:

v0.8 - master

Platforms:

Linux

Type:

int

Default:

1

Range:

0

taskq threads use the default Linux kernel thread priority

1

Change:

Dynamic

Tags:

CPU, SPL, taskq

Allow newly created taskq threads to set a non-default scheduler priority. When enabled, the priority specified when a taskq is created will be applied to all threads created by that taskq. When disabled all threads will use the default Linux kernel thread priority. By default, this behavior is enabled.

When to change: when troubleshooting CPU scheduling-related performance issues

Notes: spl_taskq_thread_priority allows newly created taskq threads to set a non-default scheduler priority. When enabled the priority specified when a taskq is created will be applied to all threads created by that taskq. When disabled all threads will use the default Linux kernel thread priority.

spl_taskq_thread_sequential

Versions:

v0.8 - master

Platforms:

Linux

Type:

uint

Default:

4

Range:

1 to MAX_INT

Change:

Dynamic

Tags:

CPU, SPL, taskq

The number of items a taskq worker thread must handle without interruption before requesting a new worker thread be spawned. This is used to control how quickly taskqs ramp up the number of threads processing the queue. Because Linux thread creation and destruction are relatively inexpensive a small default value has been selected. This means that normally threads will be created aggressively which is desirable. Increasing this value will result in a slower thread creation rate which may be preferable for some configurations.

Notes: spl_taskq_thread_sequential is the number of items a taskq worker thread must handle without interruption before requesting a new worker thread be spawned. spl_taskq_thread_sequential controls how quickly taskqs ramp up the number of threads processing the queue. Because Linux thread creation and destruction are relatively inexpensive a small default value has been selected. Thus threads are created aggressively, which is typically desirable. Increasing this value results in a slower thread creation rate which may be preferable for some configurations.

spl_taskq_thread_timeout_ms

Versions:

v2.2 - master

Platforms:

Linux

Type:

uint

Default:

5000

Change:

Dynamic

Tags:

SPL, taskq

Minimum idle threads exit interval for dynamic taskqs. Smaller values allow idle threads exit more often and potentially be respawned again on demand, causing more churn.

vdev_file_logical_ashift

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

9

Change:

Dynamic

Tags:

vdev, vdev_file

Logical ashift for file-based devices.

vdev_file_physical_ashift

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

9

Change:

Dynamic

Tags:

vdev, vdev_file

Physical ashift for file-based devices.

vdev_raidz_outlier_check_interval_ms

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1000

Change:

Dynamic

Tags:

raidz, vdev

How often each RAID-Z and dRAID vdev will check for slow disk outliers. Increasing this interval will reduce the sensitivity of detection (since all I/Os since the last check are included in the statistics), but will slow the response to a disk developing a problem. Defaults to once per second; setting extremely small values may cause negative performance effects.

vdev_raidz_outlier_insensitivity

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

50

Change:

Dynamic

Tags:

raidz, vdev

When performing slow outlier checks for RAID-Z and dRAID vdevs, this value is used to determine how far out an outlier must be before it counts as an event worth consdering. This is phrased as “insensitivity” because larger values result in fewer detections. Smaller values will result in more aggressive sitting out of disks that may have problems, but may significantly increase the rate of spurious sit-outs.

To provide a more technical definition of this parameter, this is the multiple of the inter-quartile range (IQR) that is being used in a Tukey’s Fence detection algorithm. This is much higher than a normal Tukey’s Fence k-value, because the distribution under consideration is probably an extreme-value distribution, rather than a more typical Gaussian distribution.

vdev_read_sit_out_secs

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

600

Change:

Dynamic

Tags:

raidz, vdev

When a slow disk outlier is detected it is placed in a sit out state. While sitting out the disk will not participate in normal reads, instead its data will be reconstructed as needed from parity. Scrub operations will always read from a disk, even if it’s sitting out. A number of disks in a RAID-Z or dRAID vdev may sit out at the same time, up to the number of parity devices. Writes will still be issued to a disk which is sitting out to maintain full redundancy. Defaults to 600 seconds and a value of zero disables disk sit-outs in general, including slow disk outlier detection.

vdev_removal_max_span

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

32768

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

vdev, vdev_removal

During top-level vdev removal, chunks of data are copied from the vdev which may include free space in order to trade bandwidth for IOPS. This parameter determines the maximum span of free space, in bytes, which will be included as “unnecessary” data in a chunk of copied data.

The default value here was chosen to align with zfs_vdev_read_gap_limit, which is a similar concept when doing regular reads (but there’s no reason it has to be the same).

Notes: During top-level vdev removal, chunks of data are copied from the vdev which may include free space in order to trade bandwidth for IOPS. vdev_removal_max_span sets the maximum span of free space included as unnecessary data in a chunk of copied data.

vdev_validate_skip

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

validate labels during pool import

1

do not validate vdev labels during pool import

Change:

Dynamic

Tags:

vdev

Skip label validation steps during pool import. Changing is not recommended unless you know what you’re doing and are recovering a damaged label.

When to change: do not change

Notes: vdev_validate_skip disables label validation steps during pool import. Changing is not recommended unless you know what you are doing and are recovering a damaged label.

zap_iterate_prefetch

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

prefetch, ZAP, zap_fat

If set, when we start iterating over a ZAP object, prefetch the entire object (all leaf blocks). However, this is limited by dmu_prefetch_max.

zap_micro_max_size

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

131072

Change:

Dynamic

Tags:

ZAP

Maximum micro ZAP size. A “micro” ZAP is upgraded to a “fat” ZAP once it grows beyond the specified size. Sizes higher than 128KiB will be clamped to 128KiB unless the large_microzap feature is enabled.

zap_shrink_enabled

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

ZAP, zap_fat

If set, adjacent empty ZAP blocks will be collapsed, reducing disk space.

zfetch_array_rd_sz

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

1048576

Range:

0 to MAX_ULONG

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

If prefetching is enabled, disable prefetching for reads larger than this size.

When to change: To allow prefetching when using large block sizes

Notes: If prefetching is enabled, do not prefetch blocks larger than zfetch_array_rd_sz size.

zfetch_block_cap

Versions:

v0.6

Platforms:

Linux, FreeBSD

Type:

uint

Default:

256

Change:

Dynamic

Tags:

DMU, dmu_zfetch

Max number of blocks to prefetch at a time

Default value: 256.

zfetch_hole_shift

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

2

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

Max log2 fraction of holes in a stream

zfetch_max_distance

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.1 - master

67108864

v0.7 - v2.0

8,388,608

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

Max bytes to prefetch per stream.

When to change: Consider increasing read workloads that use large blocks and exhibit high prefetch hit ratios

zfetch_max_idistance

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

67108864

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

Max bytes to prefetch indirects for per stream.

zfetch_max_reorder

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

16777216

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

Requests within this byte distance from the current prefetch stream position are considered parts of the stream, reordered due to parallel processing. Such requests do not advance the stream position immediately unless zfetch_hole_shift fill threshold is reached, but saved to fill holes in the stream later.

zfetch_max_sec_reap

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

2

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

Max time before inactive prefetch stream can be deleted

zfetch_max_streams

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

8

Range:

1 to MAX_UINT

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

Max number of streams per zfetch (prefetch streams per file).

When to change: If the workload benefits from prefetching and has more than zfetch_max_streams concurrent reader threads

Notes: Maximum number of prefetch streams per file. For version v0.7.0 and later, when prefetching small files the number of prefetch streams is automatically reduced below to prevent the streams from overlapping.

zfetch_min_distance

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

4194304

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

Min bytes to prefetch per stream. Prefetch distance starts from the demand access size and quickly grows to this value, doubling on each hit. After that it may grow further by 1/8 per hit, but only if some prefetch since last time haven’t completed in time to satisfy demand request, i.e. prefetch depth didn’t cover the read latency or the pool got saturated.

zfetch_min_sec_reap

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.1 - master

1

v0.6 - v2.0

2

Range:

0 to MAX_UINT

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

Min time before inactive prefetch stream can be reclaimed

When to change: To test prefetch efficiency

Notes: Prefetch streams that have been accessed in zfetch_min_sec_reap seconds are automatically stopped.

zfs_abd_scatter_enabled

Versions:

v0.7 - master

Platforms:

Linux

Type:

int

Default:

1

Range:

0

use linear allocation only

1

allow scatter/gather

Change:

Dynamic

Tags:

ABD, memory

Enables ARC from using scatter/gather lists and forces all allocations to be linear in kernel memory. Disabling can improve performance in some code paths at the expense of fragmented kernel memory.

When to change: Testing ABD

Verification: ABD statistics are observable in /proc/spl/kstat/zfs/abdstats. Slab allocations are observable in /proc/slabinfo

Notes: zfs_abd_scatter_enabled controls the ARC Buffer Data (ABD) scatter/gather feature. When disabled, the legacy behaviour is selected using linear buffers. For linear buffers, all the data in the ABD is stored in one contiguous buffer in memory (from a zio_[data_]buf_* kmem cache). When enabled (default), the data in the ABD is split into equal-sized chunks (from the abd_chunk_cache kmem_cache), with pointers to the chunks recorded in an array at the end of the ABD structure. This allows more efficient memory allocation for buffers, especially when large recordsizes are used.

zfs_abd_scatter_max_order

Versions:

v0.7 - master

Platforms:

Linux

Type:

uint

Default:

v2.1 - master

MAX_ORDER-1

v2.0

10 at the time of this writing

Range:

1 to 10 (upper limit is hardware-dependent)

Change:

Dynamic

Tags:

ABD, memory

Maximum number of consecutive memory pages allocated in a single block for scatter/gather lists.

The value of MAX_ORDER depends on kernel configuration.

When to change: Testing ABD features

Verification: ABD statistics are observable in /proc/spl/kstat/zfs/abdstats

Notes: zfs_abd_scatter_max_order sets the maximum order for physical page allocation when ABD is enabled (see zfs_abd_scatter_enabled) See also Buddy Memory Allocation in the Linux kernel documentation.

zfs_abd_scatter_min_size

Versions:

v0.8 - master

Platforms:

Linux

Type:

int

Default:

1536

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

ARC

This is the minimum allocation size that will use scatter (page-based) ABDs. Smaller allocations will use linear ABDs.

When to change: debugging memory allocation, especially for large pages

Notes: zfs_abd_scatter_min_size changes the ARC buffer data (ABD) allocator’s threshold for using linear or page-based scatter buffers. Allocations smaller than zfs_abd_scatter_min_size use linear ABDs. Scatter ABD’s use at least one page each, so sub-page allocations waste some space when allocated as scatter allocations. For example, 2KB scatter allocation wastes half of each page. Using linear ABD’s for small allocations results in slabs containing many allocations. This can improve memory efficiency, at the expense of more work for ARC evictions attempting to free pages, because all the buffers on one slab need to be freed in order to free the slab and its underlying pages. Typically, 512B and 1KB kmem caches have 16 buffers per slab, so it’s possible for them to actually waste more memory than scatter allocations: - one page per buf = wasting 3/4 or 7/8 - one buf per slab = wasting 15/16 Spill blocks are typically 512B and are heavily used on systems running selinux with the default dnode size and the xattr=sa property set. By default, linear allocations for 512B and 1KB, and scatter allocations for larger (>= 1.5KB) allocation requests.

zfs_active_allocator

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

charp

Default:

dynamic

Change:

Dynamic

Tags:

metaslab

Select the SPA metaslab allocator. Valid values are dynamic and cursor.

zfs_admin_snapshot

Versions:

v0.6 - master

Platforms:

Linux

Type:

int

Default:

0

Range:

0

do not allow snapshot manipulation via the filesystem

1

allow snapshot manipulation via the filesystem

Change:

Dynamic

Tags:

ctldir, filesystem, snapshot

Allow the creation, removal, or renaming of entries in the .zfs/snapshot directory to cause the creation, destruction, or renaming of snapshots. When enabled, this functionality works both locally and over NFS exports which have the no_root_squash option set.

Notes: Allow the creation, removal, or renaming of entries in the .zfs/snapshot subdirectory to cause the creation, destruction, or renaming of snapshots. When enabled this functionality works both locally and over NFS exports which have the “no_root_squash” option set.

zfs_allow_redacted_dataset_mount

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

DSL, dsl_dataset

Allow datasets received with redacted send/receive to be mounted. Normally disabled because these datasets may be missing key data.

zfs_arc_average_blocksize

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

8192

Range:

512 to 16,777,216

Change:

Prior to module load

Tags:

ARC, memory

The ARC’s buffer hash table is sized based on the assumption of an average block size of this value. This works out to roughly 1 MiB of hash table per 1 GiB of physical memory with 8-byte pointers. For configurations with a known larger average block size, this value can be increased to reduce the memory footprint.

When to change: For workloads where the known average blocksize is larger, increasing zfs_arc_average_blocksize can reduce memory usage

Notes: The ARC’s buffer hash table is sized based on the assumption of an average block size of zfs_arc_average_blocksize. The default of 8 KiB uses approximately 1 MiB of hash table per 1 GiB of physical memory with 8-byte pointers.

zfs_arc_dnode_limit

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

0

Change:

Dynamic

Tags:

ARC

When the number of bytes consumed by dnodes in the ARC exceeds this number of bytes, try to unpin some of it in response to demand for non-metadata. This value acts as a ceiling to the amount of dnode metadata, and defaults to 0, which indicates that a percent which is based on zfs_arc_dnode_limit_percent of the ARC meta buffers that may be used for dnodes.

When to change: Consider increasing if arc_prune is using excessive system time and /proc/spl/kstat/zfs/arcstats shows dnode_size is near or over arc_dnode_limit

Notes: When the number of bytes consumed by dnodes in the ARC exceeds zfs_arc_dnode_limit bytes, demand for new metadata can take from the space consumed by dnodes. The default value 0, indicates that a percent which is based on zfs_arc_dnode_limit_percent of the ARC meta buffers that may be used for dnodes.

zfs_arc_dnode_limit_percent

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

10

Range:

0 to 100

Change:

Dynamic

Tags:

ARC

Percentage that can be consumed by dnodes of ARC meta buffers.

See also zfs_arc_dnode_limit, which serves a similar purpose but has a higher priority if nonzero.

When to change: Consider increasing if arc_prune is using excessive system time and /proc/spl/kstat/zfs/arcstats shows dnode_size is near or over arc_dnode_limit

Notes: Percentage of ARC metadata space that can be used for dnodes. The value calculated for zfs_arc_dnode_limit_percent can be overridden by zfs_arc_dnode_limit.

zfs_arc_dnode_reduce_percent

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Range:

0 to 100

Change:

Dynamic

Tags:

ARC

Percentage used to size dnode prune requests. The request size is the larger of two values: zfs_arc_dnode_reduce_percent applied to the dnode count above zfs_arc_dnode_limit, or zfs_arc_dnode_reduce_percent applied to the total dnode count when non-evictable metadata exceeds 3/4 of the metadata target.

When to change: Testing dnode cache efficiency

Notes: Percentage of ARC dnodes to try to evict in response to demand for non-metadata when the number of bytes consumed by dnodes exceeds zfs_arc_dnode_limit.

zfs_arc_evict_batch_limit

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

ARC

Number ARC headers to evict per sub-list before proceeding to another sub-list. This batch-style operation prevents entire sub-lists from being evicted at once but comes at a cost of additional unlocking and locking.

When to change: Testing ARC multilist features

Notes: Number ARC headers to evict per sublist before proceeding to another sublist. This batch-style operation prevents entire sublists from being evicted at once but comes at a cost of additional unlocking and locking.

zfs_arc_evict_batches_limit

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

5

Change:

Dynamic

Tags:

ARC

Number of zfs_arc_evict_batch_limit batches to process per parallel eviction task under heavy load to reduce number of context switches.

zfs_arc_evict_threads

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Change:

Prior to module load

Tags:

ARC

Sets the number of ARC eviction threads to be used.

If set greater than 0, ZFS will dedicate up to that many threads to ARC eviction. Each thread will process one sub-list at a time, until the eviction target is reached or all sub-lists have been processed. When set to 0, ZFS will compute a reasonable number of eviction threads based on the number of CPUs.

CPUs

Threads

1-4

1

5-8

2

9-15

3

16-31

4

32-63

6

64-95

8

96-127

9

128-160

11

160-191

12

192-223

13

224-255

14

256+

16

More threads may improve the responsiveness of ZFS to memory pressure. This can be important for performance when eviction from the ARC becomes a bottleneck for reads and writes.

This parameter can only be set at module load time.

zfs_arc_eviction_pct

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

200

Change:

Dynamic

Tags:

ARC

When arc_is_overflowing(), arc_get_data_impl() waits for this percent of the requested amount of data to be evicted. For example, by default, for every 2 KiB that’s evicted, 1 KiB of it may be “reused” by a new allocation. Since this is above 100%, it ensures that progress is made towards getting arc_size under arc_c. Since this is finite, it ensures that allocations can still happen, even during the potentially long time that arc_size is more than arc_c.

zfs_arc_free_target

Versions:

v2.2 - master

Platforms:

FreeBSD

Type:

uint

Tags:

ARC, memory

Desired number of free pages below which the ARC triggers reclaim. Initialized at boot to the kernel’s vm.v_free_target value and can be adjusted at runtime. This parameter is FreeBSD-specific and uses pages, unlike the Linux-specific zfs_arc_sys_free which is measured in bytes.

When to change: When the ARC is not releasing memory fast enough to keep up with other system demands

Notes: zfs_arc_free_target is the desired number of free pages below which the ARC triggers reclaim. Initialized at boot to the kernel’s vm.v_free_target value. This parameter is FreeBSD-specific and uses pages as its unit, unlike zfs_arc_sys_free which is measured in bytes.

zfs_arc_grow_retry

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v0.7 - master

0

v0.6

5

Range:

0=use the built-in default of 5 seconds, otherwise 1 to INT_MAX

Change:

Dynamic

Tags:

ARC, memory

If set to a non zero value, it will replace the arc_grow_retry value with this value. The arc_grow_retry value (default 5s) is the number of seconds the ARC will wait before trying to resume growth after a memory pressure event.

Notes: When the ARC is shrunk due to memory demand, do not retry growing the ARC for zfs_arc_grow_retry seconds. This operates as a damper to prevent oscillating grow/shrink cycles when there is memory pressure. If zfs_arc_grow_retry = 0, the internal default of 5 seconds is used.

zfs_arc_lotsfree_percent

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

10

Range:

0 to 100

Change:

Dynamic

Tags:

ARC, memory

Throttle I/O when free system memory drops below this percentage of total system memory. Setting this value to 0 will disable the throttle.

Notes: Throttle ARC memory consumption, effectively throttling I/O, when free system memory drops below this percentage of total system memory. Setting zfs_arc_lotsfree_percent to 0 disables the throttle. The arcstat_memory_throttle_count counter in /proc/spl/kstat/arcstats can indicate throttle activity.

zfs_arc_max

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

0

Range:

67,108,864 to RAM size in bytes

Change:

Dynamic

Tags:

ARC, memory

Max size of ARC in bytes. If 0, then the max size of ARC is determined by the amount of system memory installed. The larger of all_system_memory - 1 GiB and 5/8 × all_system_memory will be used as the limit. This value must be at least 67108864B (64 MiB).

This value can be changed dynamically, with some caveats. It cannot be set back to 0 while running, and reducing it below the current ARC size will not cause the ARC to shrink without memory pressure to induce shrinking.

When to change: Reduce if ARC competes too much with other applications, increase if ZFS is the primary application and can use more RAM

Verification: c column of zarcstat (arcstat before 2.4.0) or /proc/spl/kstat/zfs/arcstats entry c_max

Notes:

If set to 0, the limit is derived from the amount of memory installed. Since v2.3.0 every platform uses the larger of all_system_memory - 1 GiB and 5/8 x all_system_memory. Before that the platforms differed:

  • Linux: half of system memory

  • FreeBSD: the larger of all_system_memory - 1 GiB and 5/8 x all_system_memory

zfs_arc_max can be changed dynamically, with some caveats. It cannot be set back to 0 while running, and reducing it below the current ARC size will not cause the ARC to shrink without memory pressure to induce shrinking.

zfs_arc_meta_adjust_restarts

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

int

Default:

4096

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

ARC

The number of restart passes to make while scanning the ARC attempting the free buffers in order to stay below the fs_arc_meta_limit. This value should not need to be tuned but is available to facilitate performance analysis.

When to change: Testing ARC metadata adjustment feature

Notes: Removed in v2.2.0 by commit a8d83e2a (“More adaptive ARC eviction”). Replaced by adaptive eviction with zfs_arc_meta_balance. The number of restart passes to make while scanning the ARC attempting the free buffers in order to stay below the zfs_arc_meta_limit.

zfs_arc_meta_balance

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

500

Change:

Dynamic

Tags:

ARC

Balance between metadata and data on ghost hits. Values above 100 increase metadata caching by proportionally reducing effect of ghost data hits on target data/metadata rate.

When to change: When the ARC is not caching enough metadata (or too much) for your workload

Notes: zfs_arc_meta_balance controls how the ARC balances metadata and data caching. When evicted metadata is re-requested (a “ghost hit”), the ARC shifts its target to cache more metadata. This parameter scales the strength of that shift relative to data ghost hits. Higher values give more preference to metadata. The default of 500 means metadata ghost hits have 5x the effect of data ghost hits. A value of 100 weights them equally. This parameter replaced the manual metadata limit tunables (zfs_arc_meta_limit, zfs_arc_meta_min, etc.) that were removed in v2.2.0.

zfs_arc_meta_limit

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

0

Range:

0 to c_max

Change:

Dynamic

Tags:

ARC

The maximum allowed size in bytes that metadata buffers are allowed to consume in the ARC. When this limit is reached, metadata buffers will be reclaimed, even if the overall arc_c_max has not been reached. It defaults to 0, which indicates that a percentage based on zfs_arc_meta_limit_percent of the ARC may be used for metadata.

This value my be changed dynamically, except that must be set to an explicit value (cannot be set back to 0).

When to change: For workloads where the metadata to data ratio in the ARC can be changed to improve ARC hit rates

Verification: /proc/spl/kstat/zfs/arcstats entry arc_meta_limit

Notes: Removed in v2.2.0 by commit a8d83e2a (“More adaptive ARC eviction”). Replaced by adaptive eviction with zfs_arc_meta_balance. Sets the maximum allowed size metadata buffers in the ARC. When zfs_arc_meta_limit is reached metadata buffers are reclaimed, even if the overall c_max has not been reached. In version v0.7.0, with a default value = 0, zfs_arc_meta_limit_percent is used to set arc_meta_limit

zfs_arc_meta_limit_percent

Versions:

v0.7 - v2.1

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

75

Range:

0 to 100

Change:

Dynamic

Tags:

ARC

Percentage of ARC buffers that can be used for metadata.

See also zfs_arc_meta_limit, which serves a similar purpose but has a higher priority if nonzero.

When to change: For workloads where the metadata to data ratio in the ARC can be changed to improve ARC hit rates

Verification: /proc/spl/kstat/zfs/arcstats entry arc_meta_limit

Notes: Removed in v2.2.0 by commit a8d83e2a (“More adaptive ARC eviction”). Replaced by adaptive eviction with zfs_arc_meta_balance. Sets the limit to ARC metadata, arc_meta_limit, as a percentage of the maximum size target of the ARC, c_max Prior to version v0.7.0, the zfs_arc_meta_limit was used to set the limit as a fixed size. zfs_arc_meta_limit_percent provides a more convenient interface for setting the limit.

zfs_arc_meta_min

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

0

Range:

16,777,216 to c_max

Change:

Dynamic

Tags:

ARC

The minimum allowed size in bytes that metadata buffers may consume in the ARC.

Verification: /proc/spl/kstat/zfs/arcstats entry arc_meta_min

Notes: Removed in v2.2.0 by commit a8d83e2a (“More adaptive ARC eviction”). Replaced by adaptive eviction with zfs_arc_meta_balance. The minimum allowed size in bytes that metadata buffers may consume in the ARC. This value defaults to 0 which disables a floor on the amount of the ARC devoted meta data. When evicting data from the ARC, if the metadata_size is less than arc_meta_min then data is evicted instead of metadata.

zfs_arc_meta_prune

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

int

Default:

10000

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

ARC

The number of dentries and inodes to be scanned looking for entries which can be dropped. This may be required when the ARC reaches the zfs_arc_meta_limit because dentries and inodes can pin buffers in the ARC. Increasing this value will cause to dentry and inode caches to be pruned more aggressively. Setting this value to 0 will disable pruning the inode and dentry caches.

Notes: Removed in v2.2.0 by commit a8d83e2a (“More adaptive ARC eviction”). Replaced by adaptive eviction with zfs_arc_meta_balance. zfs_arc_meta_prune sets the number of dentries and znodes to be scanned looking for entries which can be dropped. This provides a mechanism to ensure the ARC can honor the arc_meta_limit and reclaim otherwise pinned ARC buffers. Pruning may be required when the ARC size drops to arc_meta_limit because dentries and znodes can pin buffers in the ARC. Increasing this value will cause to dentry and znode caches to be pruned more aggressively and the arc_prune thread becomes more active. Setting zfs_arc_meta_prune to 0 will disable pruning.

zfs_arc_meta_strategy

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

evict metadata only

1

also evict data buffers if they can free metadata buffers for eviction

Change:

Dynamic

Tags:

ARC

Define the strategy for ARC metadata buffer eviction (meta reclaim strategy):

0 (META_ONLY)

evict only the ARC metadata buffers

1 (BALANCED)

additional data buffers may be evicted if required to evict the required number of metadata buffers.

When to change: Testing ARC metadata eviction

Notes: Removed in v2.2.0 by commit a8d83e2a (“More adaptive ARC eviction”). Replaced by adaptive eviction with zfs_arc_meta_balance. Defines the strategy for ARC metadata eviction (meta reclaim strategy). A value of 0 (META_ONLY) will evict only the ARC metadata. A value of 1 (BALANCED) indicates that additional data may be evicted if required in order to evict the requested amount of metadata.

zfs_arc_min

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v0.7 - master

0

v0.6

100

Range:

33,554,432 to c_max

Change:

Dynamic

Tags:

ARC

Min size of ARC in bytes. If set to 0, arc_c_min will default to consuming the larger of 32 MiB and all_system_memory / 32.

When to change: If the primary focus of the system is ZFS, then increasing can ensure the ARC gets a minimum amount of RAM

Verification: /proc/spl/kstat/zfs/arcstats entry c_min

Notes: Minimum ARC size limit. When the ARC is asked to shrink, it will stop shrinking at c_min as tuned by zfs_arc_min.

zfs_arc_min_prefetch_lifespan

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Default:

v0.7

0

v0.6

100

Range:

0 = use default value

Change:

Dynamic

Tags:

ARC, prefetch

Minimum time prefetched blocks are locked in the ARC, specified in jiffies. A value of 0 will default to 1 second.

Default value: 0.

Notes: Removed in v0.8.0 by commit d4a72f238 (“Sequential scrub and resilvers”). Prefetch lifetime logic was reworked as part of the sequential scan changes. Replaced by zfs_arc_min_prefetch_ms and zfs_arc_min_prescient_prefetch_ms. arc_min_prefetch_lifespan is the minimum time for a prefetched block to remain in ARC before it is eligible for eviction.

zfs_arc_min_prefetch_ms

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Units:

ms

Range:

0=use the built-in default of 1 second, otherwise 1 to INT_MAX

Change:

Dynamic

Tags:

ARC, prefetch

Minimum time prefetched blocks are locked in the ARC.

Notes: Minimum time prefetched blocks are locked in the ARC. A value of 0 represents the default of 1 second. However, once changed, dynamically setting to 0 will not return to the default.

zfs_arc_min_prescient_prefetch_ms

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Units:

ms

Range:

0=use the built-in default of 6 seconds, otherwise 1 to INT_MAX

Change:

Dynamic

Tags:

ARC, prefetch

Minimum time “prescient prefetched” blocks are locked in the ARC. These blocks are meant to be prefetched fairly aggressively ahead of the code that may use them.

Notes: Minimum time “prescient prefetched” blocks are locked in the ARC. These blocks are meant to be prefetched fairly aggressively ahead of the code that may use them. A value of 0 represents the default of 6 seconds. However, once changed, dynamically setting to 0 will not return to the default.

zfs_arc_no_grow_shift

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

5

Change:

Dynamic

Tags:

ARC

If less than arc_c >> zfs_arc_no_grow_shift free memory is available, the ARC is not allowed to grow.

zfs_arc_num_sublists_per_state

Versions:

v0.6

Platforms:

Linux, FreeBSD

Type:

int

Default:

1 or the number of on-online CPUs, whichever is greater

Change:

Dynamic

Tags:

ARC

To allow more fine-grained locking, each ARC state contains a series of lists for both data and meta data objects. Locking is performed at the level of these “sub-lists”. This parameters controls the number of sub-lists per ARC state.

Default value: 1 or the number of on-online CPUs, whichever is greater

zfs_arc_p_aggressive_disable

Versions:

v0.6

Platforms:

Linux, FreeBSD

Type:

int

Change:

Dynamic

Tags:

ARC

Disable aggressive arc_p growth

Use 1 for yes (default) and 0 to disable.

zfs_arc_p_dampener_disable

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

avoid large adjustments

1

permit large adjustments

Change:

Dynamic

Tags:

ARC

Disable arc_p adapt dampener, which reduces the maximum single adjustment to arc_p.

When to change: Testing ARC ghost list behaviour

Notes: Removed in v2.2.0 by commit a8d83e2a (“More adaptive ARC eviction”). Replaced by adaptive eviction with zfs_arc_meta_balance. When data is being added to the ghost lists, the MRU target size is adjusted. The amount of adjustment is based on the ratio of the MRU/MFU sizes. When enabled, the ratio is capped to 10, avoiding large adjustments.

zfs_arc_p_min_shift

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

int

Default:

v0.7 - v2.1

0

v0.6

4

Range:

0=use the built-in default of 4, otherwise 1 to INT_MAX

Change:

Dynamic

Tags:

ARC

If nonzero, this will update arc_p_min_shift (default 4) with the new value. arc_p_min_shift is used as a shift of arc_c when calculating the minumum arc_p size.

Verification: Observe changes to /proc/spl/kstat/zfs/arcstats entry p

Notes: Removed in v2.2.0 by commit a8d83e2a (“More adaptive ARC eviction”). Replaced by adaptive eviction with zfs_arc_meta_balance. arc_p_min_shift is used to shift of ARC target size (/proc/spl/kstat/zfs/arcstats entry c) for calculating both minimum and maximum most recently used (MRU) target size (/proc/spl/kstat/zfs/arcstats entry p) A value of 0 represents the default setting of arc_p_min_shift = 4. However, once changed, dynamically setting zfs_arc_p_min_shift to 0 will not return to the default.

zfs_arc_pc_percent

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0 to 100

Change:

Dynamic

Tags:

ARC, memory

Percent of pagecache to reclaim ARC to.

This tunable allows the ZFS ARC to play more nicely with the kernel’s LRU pagecache. It can guarantee that the ARC size won’t collapse under scanning pressure on the pagecache, yet still allows the ARC to be reclaimed down to zfs_arc_min if necessary. This value is specified as percent of pagecache size (as measured by NR_ACTIVE_FILE + NR_INACTIVE_FILE), where that percent may exceed 100. This only operates during memory pressure/reclaim.

When to change: When using file systems under memory shortfall, if the page scanner causes the ARC to shrink too fast, then adjusting zfs_arc_pc_percent can reduce the shrink rate

Notes: zfs_arc_pc_percent allows ZFS arc to play more nicely with the kernel’s LRU pagecache. It can guarantee that the arc size won’t collapse under scanning pressure on the pagecache, yet still allows arc to be reclaimed down to zfs_arc_min if necessary. This value is specified as percent of pagecache size (as measured by NR_FILE_PAGES) where that percent may exceed 100. This only operates during memory pressure/reclaim.

zfs_arc_prune_task_threads

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Change:

Dynamic

Tags:

ARC

Number of arc_prune threads. FreeBSD does not need more than one. Linux may theoretically use one per mount point up to number of CPUs, but that was not proven to be useful.

zfs_arc_shrink_shift

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v0.7 - master

0

v0.6

5

Range:

0=use the built-in default of 7, otherwise 1 to INT_MAX

Change:

Dynamic

Tags:

ARC, memory

If nonzero, this will update arc_shrink_shift (default 7) with the new value.

When to change: During memory shortfall, reducing zfs_arc_shrink_shift increases the rate of ARC shrinkage

Notes: arc_shrink_shift is used to adjust the ARC target sizes when large reduction is required. The current ARC target size, c, and MRU size p can be reduced by by the current size >> arc_shrink_shift. For the default value of 7, this reduces the target by approximately 0.8%. A value of 0 represents the default setting of arc_shrink_shift = 7. However, once changed, dynamically setting arc_shrink_shift to 0 will not return to the default.

zfs_arc_shrinker_limit

Versions:

v2.0 - master

Platforms:

Linux

Type:

int

Default:

v2.3 - master

0

v2.0 - v2.2

10,000

Change:

Dynamic

Tags:

ARC

This is a limit on how many pages the ARC shrinker makes available for eviction in response to one page allocation attempt. Note that in practice, the kernel’s shrinker can ask us to evict up to about four times this for one allocation attempt. To reduce OOM risk, this limit is applied for kswapd reclaims only.

For example a value of 10000 (in practice, 160 MiB per allocation attempt with 4 KiB pages) limits the amount of time spent attempting to reclaim ARC memory to less than 100 ms per allocation attempt, even with a small average compressed block size of ~8 KiB.

The parameter can be set to 0 (zero) to disable the limit, and only applies on Linux.

zfs_arc_shrinker_seeks

Versions:

v2.2 - master

Platforms:

Linux

Type:

int

Default:

2

Change:

Prior to module load

Tags:

ARC

Relative cost of ARC eviction on Linux, AKA number of seeks needed to restore evicted page. Bigger values make ARC more precious and evictions smaller, comparing to other kernel subsystems. Value of 4 means parity with page cache.

zfs_arc_sys_free

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

0

Range:

0 to ULONG_MAX

Change:

Dynamic

Tags:

ARC, memory

The target number of bytes the ARC should leave as free memory on the system. If zero, equivalent to the bigger of 512 KiB and all_system_memory/64.

When to change: Change if more free memory is desired as a margin against memory demand by applications

Notes: zfs_arc_sys_free is the target number of bytes the ARC should leave as free memory on the system. Defaults to the larger of 1/64 of physical memory or 512K. Setting this option to a non-zero value will override the default. However, once changed, dynamically setting zfs_arc_sys_free to 0 will not return to the default.

zfs_async_block_max_blocks

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.2 - master

UINT64_MAX

v2.0 - v2.1

ULONG_MAX (unlimited)

v0.8

100,000

Range:

1 to MAX_ULONG

Change:

Dynamic

Tags:

delete, DMU, DSL, dsl_scan

Maximum number of blocks freed in a single TXG.

Notes: zfs_async_block_max_blocks limits the number of blocks freed in a single transaction group commit. During deletes of large objects, such as snapshots, the number of freed blocks can cause the DMU to extend txg sync times well beyond zfs_txg_timeout. zfs_async_block_max_blocks is used to limit these effects.

zfs_async_free_zio_wait_interval

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

2000

Change:

Dynamic

Tags:

DSL, dsl_scan

After freeing this many dedup, clone or gang blocks wait for all pending I/Os to complete before continuing.

zfs_autoimport_disable

Versions:

v0.6 - v2.3

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

read zpool.cache at module load

1

do not read zpool.cache at module load

Change:

Dynamic

Tags:

import, SPA, spa_config

Disable pool import at module load by ignoring the cache file (spa_config_path).

When to change: Leave as default so that zfs behaves as other Linux kernel modules

Notes: Disable reading zpool.cache file (see spa_config_path) when loading the zfs module.

zfs_bclone_enabled

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

vnops

Enables access to the block cloning feature. If this setting is 0, then even if feature@block_cloning is enabled, using functions and system calls that attempt to clone blocks will act as though the feature is disabled.

zfs_bclone_strict_properties

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

vnops

Restricts block cloning between datasets with different properties (checksum, compression, copies, dedup, or special_small_blocks).

zfs_bclone_wait_dirty

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

v2.3 - master

1

v2.2

0

Range:

1 | 0

Change:

Dynamic

Tags:

vnops

When set to 1 the FICLONE and FICLONERANGE ioctls will wait for any dirty data to be written to disk before proceeding. This ensures that the clone operation reliably succeeds, even if a file is modified and then immediately cloned. Note that for small files this may be slower than simply copying the file. When set to 0 the clone operation will immediately fail if it encounters any dirty blocks. By default waiting is enabled. A FIDEDUPERANGE request always waits, as it has no copy fallback that could absorb the failure.

zfs_btree_verify_intensity

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Change:

Dynamic

Tags:

btree

Enables btree verification. The following settings are cumulative:

Value

Description

1

Verify height.

2

Verify pointers from children to parent.

3

Verify element counts.

4

Verify element order. (expensive)

*

5

Verify unused memory is poisoned. (expensive)

* Requires debug build.

zfs_ccw_retry_interval

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Default:

300

Change:

Dynamic

Tags:

SPA

Interval, in seconds, at which a failed write of the configuration cache file is retried.

zfs_checksum_events_per_second

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Range:

zed threshold to MAX_UINT

Change:

Dynamic

Tags:

vdev

Rate limit checksum events to this many per second. Note that this should not be set below the ZED thresholds (currently 10 checksums over 10 seconds) or else the daemon may not trigger any action.

Notes: zfs_checksum_events_per_second is a rate limit for checksum events. Note that this should not be set below the zed thresholds (currently 10 checksums over 10 sec) or else zed may not trigger any action.

zfs_checksums_per_second

Versions:

v0.7

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Change:

Dynamic

Tags:

checksum, vdev, zed

Rate limit checksum events to this many per second. Note that this should not be set below the zed thresholds (currently 10 checksums over 10 sec) or else zed may not trigger any action.

Default value: 20

When to change: If processing checksum error events at a higher rate is desired

Notes: Renamed to zfs_checksum_events_per_second in v0.8.0 (commit ad796b8a3). The ZFS Event Daemon (zed) processes events from ZFS. However, it can be overwhelmed by high rates of error reports which can be generated by failing, high-performance devices. zfs_checksums_per_second limits the rate of checksum events reported to zed. Note: do not set this value lower than the SERD limit for checksum in zed. By default, checksum_N = 10 and checksum_T = 10 minutes, resulting in a practical lower limit of 1.

zfs_commit_timeout_pct

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.2 - master

10

v0.8 - v2.1

5%

Range:

1 to 100

Change:

Dynamic

Tags:

ZIL

This controls the amount of time that a ZIL block (lwb) will remain “open” when it isn’t “full”, and it has a thread waiting for it to be committed to stable storage. The timeout is scaled based on a percentage of the last lwb latency to avoid significantly impacting the latency of each individual transaction record (itx).

Notes: zfs_commit_timeout_pct controls the amount of time that a log (ZIL) write block (lwb) remains “open” when it isn’t “full” and it has a thread waiting to commit to stable storage. The timeout is scaled based on a percentage of the last lwb latency to avoid significantly impacting the latency of each individual intent log transaction (itx).

zfs_compressed_arc_enabled

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

compressed ARC disabled (legacy behaviour)

1

compress ARC data

Change:

Dynamic

Tags:

ABD, ARC, compression

Enables storing ARC buffers in their on-disk compressed form, reducing memory pressure. When disabled, buffers are decompressed before being cached.

When to change: Testing ARC compression feature

Verification: raw ARC statistics are observable in /proc/spl/kstat/zfs/arcstats and ARC hit ratios can be observed using zarcstat (arcstat before 2.4.0)

Notes: When compression is enabled for a dataset, later reads of the data can store the blocks in ARC in their on-disk, compressed state. This can increase the effective size of the ARC, as counted in blocks, and thus improve the ARC hit ratio.

zfs_condense_indirect_commit_entry_delay_ms

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

condense, delay, vdev, vdev_indirect, vdev_removal

Vdev indirection layer (used for device removal) sleeps for this many milliseconds during mapping generation. Intended for use with the test suite to throttle vdev removal speed.

When to change: do not change

Notes: During vdev removal, the vdev indirection layer sleeps for zfs_condense_indirect_commit_entry_delay_ms milliseconds during mapping generation. This parameter is used during automated testing of the ZFS code to improve test coverage.

zfs_condense_indirect_obsolete_pct

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

25

Change:

Dynamic

Tags:

condense, vdev, vdev_indirect

Minimum percent of obsolete bytes in vdev mapping required to attempt to condense (see zfs_condense_indirect_vdevs_enable). Intended for use with the test suite to facilitate triggering condensing as needed.

zfs_condense_indirect_vdevs_enable

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not save memory

1

save memory by condensing obsolete mapping after vdev removal

Change:

Dynamic

Tags:

condense, vdev, vdev_indirect, vdev_removal

Enable condensing indirect vdev mappings. When set, attempt to condense indirect vdev mappings if the mapping uses more than zfs_condense_min_mapping_bytes bytes of memory and if the obsolete space map object uses more than zfs_condense_max_obsolete_bytes bytes on-disk. The condensing process is an attempt to save memory by removing obsolete mappings.

Notes: During vdev removal, condensing process is an attempt to save memory by removing obsolete mappings. zfs_condense_indirect_vdevs_enable enables condensing indirect vdev mappings. When set, ZFS attempts to condense indirect vdev mappings if the mapping uses more than zfs_condense_min_mapping_bytes bytes of memory and if the obsolete space map object uses more than zfs_condense_max_obsolete_bytes bytes on disk.

zfs_condense_max_obsolete_bytes

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1073741824

Range:

0 to MAX_ULONG

Change:

Dynamic

Tags:

condense, vdev, vdev_indirect, vdev_removal

Only attempt to condense indirect vdev mappings if the on-disk size of the obsolete space map object is greater than this number of bytes (see zfs_condense_indirect_vdevs_enable).

When to change: no not change

Notes: After vdev removal, zfs_condense_max_obsolete_bytes sets the limit for beginning the condensing process. Condensing begins if the obsolete space map takes up more than zfs_condense_max_obsolete_bytes of space on disk (logically). The default of 1 GiB is small enough relative to a typical pool that the space consumed by the obsolete space map is minimal. See also zfs_condense_indirect_vdevs_enable

zfs_condense_min_mapping_bytes

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

131072

Range:

0 to MAX_ULONG

Change:

Dynamic

Tags:

condense, vdev, vdev_indirect, vdev_removal

Minimum size vdev mapping to attempt to condense (see zfs_condense_indirect_vdevs_enable).

When to change: do not change

Notes: After vdev removal, zfs_condense_min_mapping_bytes is the lower limit for determining when to condense the in-memory obsolete space map. The condensing process will not continue unless a minimum of zfs_condense_min_mapping_bytes of memory can be freed. See also zfs_condense_indirect_vdevs_enable

zfs_dbgmsg_enable

Versions:

v0.6 - master

Platforms:

FreeBSD

Type:

int

Default:

v2.1 - master

1

v0.6 - v2.0

0

Range:

0

do not log debug messages

1

log debug messages

Change:

Dynamic

Tags:

debug

Internally ZFS keeps a small log to facilitate debugging. The log is enabled by default, and can be disabled by unsetting this option. The contents of the log can be accessed by reading /proc/spl/kstat/zfs/dbgmsg. Writing 0 to the file clears the log.

This setting does not influence debug prints due to zfs_flags.

When to change: To view ZFS internal debug log

Notes: Internally ZFS keeps a small log to facilitate debugging. The contents of the log are in the /proc/spl/kstat/zfs/dbgmsg file. Writing 0 to /proc/spl/kstat/zfs/dbgmsg file clears the log. See also zfs_dbgmsg_maxsize

zfs_dbgmsg_maxsize

Versions:

v0.6 - master

Platforms:

FreeBSD

Type:

uint

Default:

v2.1 - master

4194304

v0.6 - v2.0

4M

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

debug

Maximum size of the internal ZFS debug log.

Notes: The /proc/spl/kstat/zfs/dbgmsg file size limit is set by zfs_dbgmsg_maxsize. See also zfs_dbgmsg_enable

zfs_dbuf_state_index

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Change:

Dynamic

Tags:

dbuf, debug

Historically used for controlling what reporting was available under /proc/spl/kstat/zfs. No effect.

When to change: Do not change

Notes: The zfs_dbuf_state_index feature is currently unused. It is normally used for controlling values in the /proc/spl/kstat/zfs/dbufs file.

zfs_ddt_data_is_special

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not use special vdevs to store DDT

1

store DDT in special vdevs

Change:

Dynamic

Tags:

dedup, SPA, special_vdev

If enabled, ZFS will place DDT data into the special allocation class.

When to change: when using a special top-level vdev and no dedup top-level vdev and it is desired to store the DDT in the main pool top-level vdevs

Notes: zfs_ddt_data_is_special enables the deduplication table (DDT) to reside on a special top-level vdev.

zfs_deadman_checktime_ms

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v0.8 - master

60,000

v0.7

5,000

Range:

1 to ULONG_MAX

Change:

Dynamic

Tags:

deadman, debug, SPA

Check time in milliseconds. This defines the frequency at which we check for hung I/O requests and potentially invoke the zfs_deadman_failmode behavior.

When to change: When debugging slow I/O

Notes: Once a pool sync operation has taken longer than zfs_deadman_synctime_ms milliseconds, continue to check for slow operations every zfs_deadman_checktime_ms milliseconds.

zfs_deadman_enabled

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not log slow I/O

1

log slow I/O

Change:

Dynamic

Tags:

deadman, debug, SPA

When a pool sync operation takes longer than zfs_deadman_synctime_ms, or when an individual I/O operation takes longer than zfs_deadman_ziotime_ms, then the operation is considered to be “hung”. If zfs_deadman_enabled is set, then the deadman behavior is invoked as described by zfs_deadman_failmode. By default, the deadman is enabled and set to wait which results in “hung” I/O operations only being logged. The deadman is automatically disabled when a pool gets suspended.

When to change: To disable logging of slow I/O

Notes: When a pool sync operation takes longer than zfs_deadman_synctime_ms milliseconds, a “slow spa_sync” message is logged to the debug log (see zfs_dbgmsg_enable). If zfs_deadman_enabled is set to 1, then all pending IO operations are also checked and if any haven’t completed within zfs_deadman_synctime_ms milliseconds, a “SLOW IO” message is logged to the debug log and a “deadman” system event (see zpool events command) with the details of the hung IO is posted.

zfs_deadman_events_per_second

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Change:

Dynamic

Tags:

vdev

Rate limit deadman zevents (which report hung I/O operations) to this many per second.

zfs_deadman_failmode

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

charp

Default:

wait

Range:

wait

wait for the “hung” I/O (default)

continue

attempt to recover from the “hung” I/O

panic

panic the system

Change:

Dynamic

Tags:

deadman, debug, SPA

Controls the failure behavior when the deadman detects a “hung” I/O operation. Valid values are:

wait

Wait for a “hung” operation to complete. For each “hung” operation a “deadman” event will be posted describing that operation.

continue

Attempt to recover from a “hung” operation by re-dispatching it to the I/O pipeline if possible.

panic

Panic the system. This can be used to facilitate automatic fail-over to a properly configured fail-over partner.

When to change: In some cluster cases, panic can be appropriate

Notes: zfs_deadman_failmode controls the behavior of the I/O deadman timer when it detects a “hung” I/O..

zfs_deadman_synctime_ms

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v0.8 - master

600,000

v0.6 - v0.7

1,000,000

Range:

1 to ULONG_MAX

Change:

Dynamic

Tags:

deadman, debug, SPA

Interval in milliseconds after which the deadman is triggered and also the interval after which a pool sync operation is considered to be “hung”. Once this limit is exceeded the deadman will be invoked every zfs_deadman_checktime_ms milliseconds until the pool sync completes.

When to change: When debugging slow I/O

Notes: The I/O deadman timer expiration time has two meanings 1. determines when the spa_deadman() logic should fire, indicating the txg sync has not completed in a timely manner 2. determines if an I/O is considered “hung” In version v0.8.0, any I/O that has not completed in zfs_deadman_synctime_ms is considered “hung” resulting in one of three behaviors controlled by the zfs_deadman_failmode parameter. zfs_deadman_synctime_ms takes effect if zfs_deadman_enabled = 1.

zfs_deadman_ziotime_ms

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

300000

Range:

1 to ULONG_MAX

Change:

Dynamic

Tags:

deadman, debug, SPA

Interval in milliseconds after which the deadman is triggered and an individual I/O operation is considered to be “hung”. As long as the operation remains “hung”, the deadman will be invoked every zfs_deadman_checktime_ms milliseconds until the operation completes.

When to change: Testing ABD features

Notes: When an individual I/O takes longer than zfs_deadman_ziotime_ms milliseconds, then the operation is considered to be “hung”. If zfs_deadman_enabled is set then the deadman behaviour is invoked as described by the zfs_deadman_failmode option.

zfs_dedup_log_cap

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

UINT_MAX

Change:

Dynamic

Tags:

DDT, dedup

Soft cap for the size of the current dedup log.

If the log is larger than this size, we increase the aggressiveness of the flushing to try to bring it back down to the soft cap. Setting it will reduce import times, but will reduce the efficiency of the DDT log, increasing the expected number of IOs required to flush the same amount of data.

zfs_dedup_log_flush_entries_max

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

UINT_MAX

Change:

Dynamic

Tags:

DDT, dedup

Flush at most this many entries each transaction.

Mostly used for debugging purposes.

zfs_dedup_log_flush_entries_min

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

100

Change:

Dynamic

Tags:

DDT, dedup

Flush at least this many entries each transaction.

OpenZFS will flush a fraction of the log every TXG, to keep the size proportional to the ingest rate (see zfs_dedup_log_flush_txgs). This sets the minimum for that estimate, which prevents the backlog from completely draining if the ingest rate falls. Raising it can force OpenZFS to flush more aggressively, reducing the backlog to zero more quickly, but can make it less able to back off if log flushing would compete with other IO too much.

zfs_dedup_log_flush_flow_rate_txgs

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Change:

Dynamic

Tags:

DDT, dedup

Number of transactions to use to compute the flow rate.

OpenZFS will estimate number of entries changed (ingest rate), number of entries flushed (flush rate) and time spent flushing (flush time rate) and combining these into an overall “flow rate”. It will use an exponential weighted moving average over some number of recent transactions to compute these rates. This sets the number of transactions to compute these averages over. Setting it higher can help to smooth out the flow rate in the face of spiky workloads, but will take longer for the flow rate to adjust to a sustained change in the ingress rate.

zfs_dedup_log_flush_min_time_ms

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1000

Change:

Dynamic

Tags:

DDT, dedup

Minimum time to spend on dedup log flush each transaction.

At least this long will be spent flushing dedup log entries each transaction, up to zfs_txg_timeout. This occurs even if doing so would delay the transaction, that is, other IO completes under this time.

zfs_dedup_log_flush_txgs

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

100

Change:

Dynamic

Tags:

DDT, dedup

Target number of TXGs to process the whole dedup log.

Every TXG, OpenZFS will process the inverse of this number times the size of the DDT backlog. This will keep the backlog at a size roughly equal to the ingest rate times this value. This offers a balance between a more efficient DDT log, with better aggregation, and shorter import times, which increase as the size of the DDT log increases. Increasing this value will result in a more efficient DDT log, but longer import times.

zfs_dedup_log_hard_cap

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

DDT, dedup

Whether to treat the log cap as a firm cap or not.

When set to 0 (the default), the zfs_dedup_log_cap will increase the maximum number of log entries we flush in a given txg. This will bring the backlog size down towards the cap, but not at the expense of making TXG syncs take longer. If this is set to 1, the cap acts more like a hard cap than a soft cap; it will also increase the minimum number of log entries we flush per TXG. Enabling it will reduce worst-case import times, at the cost of increased TXG sync times.

zfs_dedup_log_mem_max

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

0

Change:

Prior to module load

Tags:

DDT, ddt_log, dedup

Max memory to use for dedup logs.

OpenZFS will spend no more than this much memory on maintaining the in-memory dedup log. Flushing will begin when around half this amount is being spent on logs. The default value of 0 will cause it to be set by zfs_dedup_log_mem_max_percent instead.

zfs_dedup_log_mem_max_percent

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Change:

Prior to module load

Tags:

DDT, ddt_log, dedup

Max memory to use for dedup logs, as a percentage of total memory.

If zfs_dedup_log_mem_max is not set, it will be initialized as a percentage of the total memory in the system.

zfs_dedup_log_txg_max

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

8

Change:

Dynamic

Tags:

DDT, ddt_log, dedup

Max transactions to before starting to flush dedup logs.

OpenZFS maintains two dedup logs, one receiving new changes, one flushing. If there is nothing to flush, it will accumulate changes for no more than this many transactions before switching the logs and starting to flush entries out.

zfs_dedup_prefetch

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

do not prefetch

1

prefetch dedup table entries

Change:

Dynamic

Tags:

DDT, dedup, memory, prefetch

Enable prefetching dedup-ed blocks which are going to be freed.

When to change: For systems with limited RAM using the dedup feature, disabling deduplication table prefetch can reduce memory pressure

Notes: ZFS can prefetch deduplication table (DDT) entries. zfs_dedup_prefetch allows DDT prefetches to be enabled.

zfs_default_bs

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

9

Change:

Dynamic

Tags:

dnode

Default dnode block size as a power of 2.

zfs_default_ibs

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

17

Change:

Dynamic

Tags:

dnode

Default dnode indirect block size as a power of 2.

zfs_delay_min_dirty_percent

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

60

Range:

0 to 100

Change:

Dynamic

Tags:

delay, DSL, dsl_pool, write_throttle

Start to delay each transaction once there is this amount of dirty data, expressed as a percentage of zfs_dirty_data_max. This value should be at least zfs_vdev_async_write_active_max_dirty_percent. See ZFS TRANSACTION DELAY.

When to change: See ZFS Transaction Delay

Notes: The ZFS write throttle begins to delay each transaction when the amount of dirty data reaches the threshold zfs_delay_min_dirty_percent of zfs_dirty_data_max. This value should be >= zfs_vdev_async_write_active_max_dirty_percent.

zfs_delay_scale

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

500000

Range:

0 to ULONG_MAX

Change:

Dynamic

Tags:

delay, DSL, dsl_pool, write_throttle

This controls how quickly the transaction delay approaches infinity. Larger values cause longer delays for a given amount of dirty data.

For the smoothest delay, this value should be about 1 billion divided by the maximum number of operations per second. This will smoothly handle between ten times and a tenth of this number. See ZFS TRANSACTION DELAY.

zfs_delay_scale × zfs_dirty_data_max must be smaller than 2^64.

When to change: See ZFS Transaction Delay

Notes: zfs_delay_scale controls how quickly the ZFS write throttle transaction delay approaches infinity. Larger values cause longer delays for a given amount of dirty data. For the smoothest delay, this value should be about 1 billion divided by the maximum number of write operations per second the pool can sustain. The throttle will smoothly handle between 10x and 1/10th zfs_delay_scale. Note: zfs_delay_scale * zfs_dirty_data_max must be < 2^64.

zfs_delays_per_second

Versions:

v0.7

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Change:

Dynamic

Tags:

delay, vdev, zed

Rate limit IO delay events to this many per second.

Default value: 20

When to change: If processing delay events at a higher rate is desired

Notes: Renamed to zfs_slow_io_events_per_second in v0.8.0 (commit ad796b8a3). The ZFS Event Daemon (zed) processes events from ZFS. However, it can be overwhelmed by high rates of error reports which can be generated by failing, high-performance devices. zfs_delays_per_second limits the rate of delay events reported to zed.

zfs_delete_blocks

Versions:

v0.7 - master

Platforms:

Linux

Type:

ulong

Default:

20480

Range:

1 to ULONG_MAX

Change:

Dynamic

Tags:

delete, filesystem, vnops

This is the used to define a large file for the purposes of deletion. Files containing more than zfs_delete_blocks will be deleted asynchronously, while smaller files are deleted synchronously. Decreasing this value will reduce the time spent in an unlink(2) system call, at the expense of a longer delay before the freed space is available. This only applies on Linux.

When to change: If applications delete large files and blocking on unlink(2) is not desired

Notes: zfs_delete_blocks defines a large file for the purposes of delete. Files containing more than zfs_delete_blocks will be deleted asynchronously while smaller files are deleted synchronously. Decreasing this value reduces the time spent in an unlink(2) system call at the expense of a longer delay before the freed space is available. The zfs_delete_blocks value is specified in blocks, not bytes. The size of blocks can vary and is ultimately limited by the filesystem’s recordsize property.

zfs_delete_dentry

Versions:

v2.3 - master

Platforms:

Linux

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

delete, super

Sets whether the kernel should free a dentry structure when it is no longer required, or hold it in the dentry cache. Intended for testing/debugging. Since a dentry structure holds an inode reference, a cached dentry can “pin” an inode in memory indefinitely, along with associated OpenZFS structures (See zfs_delete_inode).

The default value of 0 instructs the kernel to cache entries and their associated inodes when they are no longer directly referenced. They will be reclaimed as part of the kernel’s normal cache management processes. Setting it to 1 will instruct the kernel to release directory entries and their inodes as soon as they are no longer referenced by the filesystem.

This parameter is only available on Linux.

zfs_delete_inode

Versions:

v2.3 - master

Platforms:

Linux

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

delete, super

Sets whether the kernel should free an inode structure when the last reference is released, or cache it in memory. Intended for testing/debugging.

A live inode structure “pins” versious internal OpenZFS structures in memory, which can result in large amounts of “unusable” memory on systems with lots of infrequently-accessed files, until the kernel’s memory pressure mechanism asks OpenZFS to release them.

The default value of 0 always caches inodes that appear to still exist on disk. Setting it to 1 will immediately release unused inodes and their associated memory back to the dbuf cache or the ARC for reuse, but may reduce performance if inodes are frequently evicted and reloaded.

This parameter is only available on Linux.

zfs_dio_enabled

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

v2.4 - master

1

v2.3

0

Range:

1 | 0

Change:

Dynamic

Tags:

vnops

Enable Direct I/O. If this setting is 0, then all I/O requests will be directed through the ARC acting as though the dataset property direct was set to disabled.

zfs_dio_strict

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

vnops

Strictly enforce alignment for Direct I/O requests, returning EINVAL if not page-aligned instead of silently falling back to uncached I/O.

zfs_dio_write_verify_events_per_second

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Change:

Dynamic

Tags:

vdev

Rate limit Direct I/O write verify events to this many per second.

zfs_dirty_data_max

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Range:

1 to zfs_dirty_data_max_max

Change:

Dynamic

Tags:

DSL, dsl_pool, write_throttle

Determines the dirty space limit in bytes. Once this limit is exceeded, new writes are halted until space frees up. This parameter takes precedence over zfs_dirty_data_max_percent. See ZFS TRANSACTION DELAY.

Defaults to physical_ram/10, capped at zfs_dirty_data_max_max.

When to change: See ZFS Transaction Delay

Notes: zfs_dirty_data_max is the ZFS write throttle dirty space limit. Once this limit is exceeded, new writes are delayed until space is freed by writes being committed to the pool. zfs_dirty_data_max takes precedence over zfs_dirty_data_max_percent.

zfs_dirty_data_max_max

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

u64

Range:

1 to physical RAM size

Change:

Prior to module load

Tags:

DSL, dsl_pool, write_throttle

Maximum allowable value of zfs_dirty_data_max, expressed in bytes. This limit is only enforced at module load time, and will be ignored if zfs_dirty_data_max is later changed. This parameter takes precedence over zfs_dirty_data_max_max_percent. See ZFS TRANSACTION DELAY.

Defaults to min(physical_ram/4, 4GiB), or min(physical_ram/4, 1GiB) for 32-bit systems.

When to change: See ZFS Transaction Delay

Notes: zfs_dirty_data_max_max is the maximum allowable value of zfs_dirty_data_max. zfs_dirty_data_max_max takes precedence over zfs_dirty_data_max_max_percent.

zfs_dirty_data_max_max_percent

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

25

Range:

1 to 100

Change:

Prior to module load

Tags:

DSL, dsl_pool, write_throttle

Maximum allowable value of zfs_dirty_data_max, expressed as a percentage of physical RAM. This limit is only enforced at module load time, and will be ignored if zfs_dirty_data_max is later changed. The parameter zfs_dirty_data_max_max takes precedence over this one. See ZFS TRANSACTION DELAY.

When to change: See ZFS Transaction Delay

Notes: zfs_dirty_data_max_max_percent an alternative to zfs_dirty_data_max_max for setting the maximum allowable value of zfs_dirty_data_max zfs_dirty_data_max_max takes precedence over zfs_dirty_data_max_max_percent

zfs_dirty_data_max_percent

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Range:

1 to 100

Change:

Prior to module load

Tags:

DSL, dsl_pool, write_throttle

Determines the dirty space limit, expressed as a percentage of all memory. Once this limit is exceeded, new writes are halted until space frees up. The parameter zfs_dirty_data_max takes precedence over this one. See ZFS TRANSACTION DELAY.

Subject to zfs_dirty_data_max_max.

When to change: See ZFS Transaction Delay

Notes: zfs_dirty_data_max_percent is an alternative method of specifying zfs_dirty_data_max, the ZFS write throttle dirty space limit. Once this limit is exceeded, new writes are delayed until space is freed by writes being committed to the pool. zfs_dirty_data_max takes precedence over zfs_dirty_data_max_percent.

zfs_dirty_data_sync

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

67,108,864

Range:

1 to ULONG_MAX

Change:

Dynamic

Tags:

DSL, dsl_pool, write_throttle, ZIO_scheduler

Start syncing out a transaction group if there is at least this much dirty data.

Default value: 67,108,864.

Notes: When there is at least zfs_dirty_data_sync dirty data, a transaction group sync is started. This allows a transaction group sync to occur more frequently than the transaction group timeout interval (see zfs_txg_timeout) when there is dirty data to be written.

zfs_dirty_data_sync_percent

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Range:

1 to zfs_vdev_async_write_active_min_dirty_percent

Change:

Dynamic

Tags:

DSL, dsl_pool, write_throttle, ZIO_scheduler

Start syncing out a transaction group if there’s at least this much dirty data (as a percentage of zfs_dirty_data_max). This should be less than zfs_vdev_async_write_active_min_dirty_percent.

Notes: When there is at least zfs_dirty_data_sync_percent of zfs_dirty_data_max dirty data, a transaction group sync is started. This allows a transaction group sync to occur more frequently than the transaction group timeout interval (see zfs_txg_timeout) when there is dirty data to be written.

zfs_disable_dup_eviction

Versions:

v0.6

Platforms:

Linux, FreeBSD

Type:

int

Range:

0

duplicate buffers can be evicted

1

do not evict duplicate buffers

Change:

Dynamic

Tags:

ARC, dedup

Disable duplicate buffer eviction

Use 1 for yes and 0 for no (default).

Notes: Removed in v0.7.0 by commit d3c2ae1c0 (“OpenZFS 6950 - ARC should cache compressed data”). Duplicate buffer handling was reworked as part of the compressed ARC changes. No replacement parameter. Disable duplicate buffer eviction from ARC.

zfs_disable_ivset_guid_check

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

check IVset guid

1

do not check IVset guid

Change:

Dynamic

Tags:

DSL, receive

Disables requirement for IVset GUIDs to be present and match when doing a raw receive of encrypted datasets. Intended for users whose pools were created with OpenZFS pre-release versions and now have compatibility issues.

When to change: debugging pre-release ZFS raw sends

Notes: zfs_disable_ivset_guid_check disables requirement for IVset guids to be present and match when doing a raw receive of encrypted datasets. Intended for users whose pools were created with ZFS on Linux pre-release versions and now have compatibility issues. For a ZFS raw receive, from a send stream created by zfs send --raw, the crypt_keydata nvlist includes a to_ivset_guid to be set on the new snapshot. This value will override the value generated by the snapshot code. However, this value may not be present, because older implementations of the raw send code did not include this value. When zfs_disable_ivset_guid_check is enabled, the receive proceeds and a newly-generated value is used.

zfs_dmu_offset_next_sync

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not force txg sync to find holes

1

force txg sync to find holes

Change:

Dynamic

Tags:

DMU

Enable forcing TXG sync to find holes. When enabled forces ZFS to sync data when SEEK_HOLE or SEEK_DATA flags are used allowing holes in a file to be accurately reported. When disabled holes will not be reported in recently dirtied files.

When to change: to exchange strict hole reporting for performance

Notes: zfs_dmu_offset_next_sync enables forcing txg sync to find holes. This causes ZFS to act like older versions when SEEK_HOLE or SEEK_DATA flags are used: when a dirty dnode causes txgs to be synced so the previous data can be found.

zfs_embedded_slog_min_ms

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

64

Change:

Dynamic

Tags:

vdev

Usually, one metaslab from each normal and special class vdev is dedicated for use by the ZIL to log synchronous writes. However, if there are fewer than zfs_embedded_slog_min_ms metaslabs in the vdev, this functionality is disabled. This ensures that we don’t set aside an unreasonable amount of space for the ZIL.

zfs_expire_snapshot

Versions:

v0.6 - master

Platforms:

Linux

Type:

int

Default:

300

Range:

0 disables automatic unmounting, maximum time is INT_MAX

Change:

Dynamic

Tags:

ctldir, filesystem, snapshot

Time before expiring .zfs/snapshot.

Notes: Snapshots of filesystems are normally automounted under the filesystem’s .zfs/snapshot subdirectory. When not in use, snapshots are unmounted after zfs_expire_snapshot seconds.

zfs_fallocate_reserve_percent

Versions:

v2.0 - master

Platforms:

Linux

Type:

uint

Default:

110

Change:

Dynamic

Tags:

file

Since ZFS is a copy-on-write filesystem with snapshots, blocks cannot be preallocated for a file in order to guarantee that later writes will not run out of space. Instead, fallocate(2) space preallocation only checks that sufficient space is currently available in the pool or the user’s project quota allocation, and then creates a sparse file of the requested size. The requested space is multiplied by zfs_fallocate_reserve_percent to allow additional space for indirect blocks and other internal metadata. Setting this to 0 disables support for fallocate(2) and causes it to return EOPNOTSUPP.

zfs_flags

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Change:

Dynamic

Tags:

debug, SPA

Set additional debugging flags. The following flags may be bitwise-ored together:

Value

Name

Description

1

ZFS_DEBUG_DPRINTF

Enable dprintf entries in the debug log.

*

2

ZFS_DEBUG_DBUF_VERIFY

Enable extra dbuf verifications.

*

4

ZFS_DEBUG_DNODE_VERIFY

Enable extra dnode verifications.

8

ZFS_DEBUG_SNAPNAMES

Enable snapshot name verification.

*

16

ZFS_DEBUG_MODIFY

Check for illegally modified ARC buffers.

64

ZFS_DEBUG_ZIO_FREE

Enable verification of block frees.

128

ZFS_DEBUG_HISTOGRAM_VERIFY

Enable extra spacemap histogram verifications.

256

ZFS_DEBUG_METASLAB_VERIFY

Verify space accounting on disk matches in-memory range_trees.

512

ZFS_DEBUG_SET_ERROR

Enable SET_ERROR and dprintf entries in the debug log.

1024

ZFS_DEBUG_INDIRECT_REMAP

Verify split blocks created by device removal.

2048

ZFS_DEBUG_TRIM

Verify TRIM ranges are always within the allocatable range tree.

4096

ZFS_DEBUG_LOG_SPACEMAP

Verify that the log summary is consistent with the spacemap log

and enable zfs_dbgmsgs for metaslab loading and flushing.

8192

ZFS_DEBUG_METASLAB_ALLOC

Enable debugging messages when allocations fail.

16384

ZFS_DEBUG_BRT

Enable BRT-related debugging messages.

32768

ZFS_DEBUG_RAIDZ_RECONSTRUCT

Enabled debugging messages for raidz reconstruction.

65536

ZFS_DEBUG_DDT

Enable DDT-related debugging messages.

* Requires debug build.

When to change: When debugging ZFS

zfs_fletcher_4_impl

Versions:

v0.7 - v0.8

Platforms:

Linux, FreeBSD

Type:

string

Default:

fastest

Range:

fastest, scalar, superscalar, superscalar4, sse2, ssse3, avx2, avx512f, or aarch64_neon depending on hardware support

Change:

Dynamic

Tags:

checksum, CPU

Select a fletcher 4 implementation.

Supported selectors are: fastest, scalar, sse2, ssse3, avx2, avx512f, and aarch64_neon. All of the selectors except fastest and scalar require instruction set extensions to be available and will only appear if ZFS detects that they are present at runtime. If multiple implementations of fletcher 4 are available, the fastest will be chosen using a micro benchmark. Selecting scalar results in the original, CPU based calculation, being used. Selecting any option other than fastest and scalar results in vector instructions from the respective CPU instruction set being used.

Default value: fastest.

When to change: Testing Fletcher-4 algorithms

Notes: Fletcher-4 is the default checksum algorithm for metadata and data. When the zfs kernel module is loaded, a set of microbenchmarks are run to determine the fastest algorithm for the current hardware. The zfs_fletcher_4_impl parameter allows a specific implementation to be specified other than the default (fastest). Selectors other than fastest and scalar require instruction set extensions to be available and will only appear if ZFS detects their presence. The scalar implementation works on all processors. The results of the microbenchmark are visible in the /proc/spl/kstat/zfs/fletcher_4_bench file. Larger numbers indicate better performance. Since ZFS is processor endian-independent, the microbenchmark is run against both big and little-endian transformation.

zfs_free_bpobj_enabled

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not process free_bpobj objects

1

process free_bpobj objects

Change:

Dynamic

Tags:

delete, DSL, dsl_scan

Enable/disable the processing of the free_bpobj object.

When to change: If there’s a problem with processing free_bpobj (e.g. i/o error or bug)

zfs_free_leak_on_eio

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

normal behavior

1

ignore error and permanently leak space

Change:

Dynamic

Tags:

debug, SPA

If destroy encounters an EIO while reading metadata (e.g. indirect blocks), space referenced by the missing metadata can not be freed. Normally this causes the background destroy to become “stalled”, as it is unable to make forward progress. While in this stalled state, all remaining space to free from the error-encountering filesystem is “temporarily leaked”. Set this flag to cause it to ignore the EIO, permanently leak the space from indirect blocks that can not be read, and continue to free everything else that it can.

The default “stalling” behavior is useful if the storage partially fails (i.e. some but not all I/O operations fail), and then later recovers. In this case, we will be able to continue pool operations while it is partially failed, and when it recovers, we can continue to free the space, with no leaks. Note, however, that this case is actually fairly rare.

Typically pools either

  1. fail completely (but perhaps temporarily, e.g. due to a top-level vdev going offline), or

  2. have localized, permanent errors (e.g. disk returns the wrong data due to bit flip or firmware bug).

In the former case, this setting does not matter because the pool will be suspended and the sync thread will not be able to make forward progress regardless. In the latter, because the error is permanent, the best we can do is leak the minimum amount of space, which is what setting this flag will do. It is therefore reasonable for this flag to normally be set, but we chose the more conservative approach of not setting it, so that there is no possibility of leaking space in the “partial temporary” failure case.

When to change: When debugging I/O errors during destroy

Notes: If destroy encounters an I/O error (EIO) while reading metadata (eg indirect blocks), space referenced by the missing metadata cannot be freed. Normally, this causes the background destroy to become “stalled”, as the destroy is unable to make forward progress. While in this stalled state, all remaining space to free from the error-encountering filesystem is temporarily leaked. Set zfs_free_leak_on_eio = 1 to ignore the EIO, permanently leak the space from indirect blocks that can not be read, and continue to free everything else that it can. The default, stalling behavior is useful if the storage partially fails (eg some but not all I/Os fail), and then later recovers. In this case, we will be able to continue pool operations while it is partially failed, and when it recovers, we can continue to free the space, with no leaks. However, note that this case is rare. Typically pools either: 1. fail completely (but perhaps temporarily (eg a top-level vdev going offline) 2. have localized, permanent errors (eg disk returns the wrong data due to bit flip or firmware bug) In case (1), the zfs_free_leak_on_eio setting does not matter because the pool will be suspended and the sync thread will not be able to make forward progress. In case (2), because the error is permanent, the best effort do is leak the minimum amount of space. Therefore, it is reasonable for zfs_free_leak_on_eio be set, but by default the more conservative approach is taken, so that there is no possibility of leaking space in the “partial temporary” failure case.

zfs_free_max_blocks

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

100,000

Range:

1 to ULONG_MAX

Change:

Dynamic

Tags:

delete, DSL, dsl_scan, filesystem

Maximum number of blocks freed in a single txg.

Default value: 100,000.

When to change: For workloads that delete large files, zfs_free_max_blocks can be adjusted to meet performance requirements while reducing the impacts of deletion

Notes: zfs_free_max_blocks sets the maximum number of blocks to be freed in a single transaction group (txg). For workloads that delete (free) large numbers of blocks in a short period of time, the processing of the frees can negatively impact other operations, including txg commits. zfs_free_max_blocks acts as a limit to reduce the impact.

zfs_free_min_time_ms

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.4 - master

500

v0.6 - v2.3

1,000

Range:

1 to (zfs_txg_timeout * 1000)

Change:

Dynamic

Tags:

delete, DSL, dsl_scan

During a zfs destroy operation using the async_destroy feature, a minimum of this much time will be spent working on freeing blocks per TXG.

Notes: During a zfs destroy operation using feature@async_destroy a minimum of zfs_free_min_time_ms time will be spent working on freeing blocks per txg commit.

zfs_history_output_max

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1048576

Change:

Dynamic

Tags:

ioctl

When attempting to log an output nvlist of an ioctl in the on-disk history, the output will not be stored if it is larger than this size (in bytes). This must be less than DMU_MAX_ACCESS (64 MiB). This applies primarily to zfs_ioc_channel_program() (cf. zfs-program(8)).

zfs_immediate_write_sz

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

32768

Range:

512 to 16,777,216 (valid block sizes)

Change:

Dynamic

Tags:

ZIL

Largest write size to store the data directly into the ZIL if logbias=latency. Larger writes may be written indirectly similar to logbias=throughput. In presence of SLOG this parameter is ignored, as if it was set to infinity, storing all written data into ZIL to not depend on regular vdev latency.

Verification: Data blocks that exceed zfs_immediate_write_sz or are written as logbias=throughput increment the zil_itx_indirect_count entry in /proc/spl/kstat/zfs/zil

Notes: If a pool does not have a log device, data blocks equal to or larger than zfs_immediate_write_sz are treated as if the dataset being written to had the property setting logbias=throughput Terminology note: logbias=throughput writes the blocks in “indirect mode” to the ZIL where the data is written to the pool and a pointer to the data is written to the ZIL.

zfs_import_defer_txgs

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

5

Change:

Dynamic

Tags:

DSL, dsl_scan

Number of transaction groups to wait after pool import before starting background work such as asynchronous block freeing (from snapshots, clones, and deduplication) and scrub or resilver operations. This allows the pool import and filesystem mounting to complete more quickly without interference from background activities. The default value of 5 transaction groups typically provides sufficient time for import and mount operations to complete on most systems.

zfs_initialize_chunk_size

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1048576

Change:

Dynamic

Tags:

vdev

Size of writes used by zpool-initialize(8). This option is used by the test suite.

zfs_initialize_value

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

16045690984833335022

Change:

Dynamic

Tags:

vdev, vdev_initialize

Default pattern written to vdev free space by zpool-initialize(8), used unless zpool initialize -z requests zeroes instead. The value in effect when a device starts initializing is recorded per device and reused if the run is suspended and later resumed.

When to change: when debugging initialization code

Notes: When initializing a vdev, ZFS writes patterns of zfs_initialize_value bytes to the device.

zfs_keep_log_spacemaps_at_export

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

Prevent log spacemaps from being destroyed during pool exports and destroys.

zfs_key_max_salt_uses

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

400000000

Range:

1 to ULONG_MAX

Change:

Dynamic

Tags:

encryption, ZIO

Maximum number of uses of a single salt value before generating a new one for encrypted datasets. The default value is also the maximum.

When to change: Testing encryption features

Notes: For encrypted datasets, the salt is regenerated every zfs_key_max_salt_uses blocks. This automatic regeneration reduces the probability of collisions due to the Birthday problem. When set to the default (400,000,000) the probability of collision is approximately 1 in 1 trillion.

zfs_livelist_condense_new_alloc

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Change:

Dynamic

Tags:

livelist, livelist_condense, SPA

Incremented each time an extra ALLOC blkptr is added to a livelist entry while it is being condensed. This option is used by the test suite to track race conditions.

zfs_livelist_condense_sync_cancel

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Change:

Dynamic

Tags:

livelist, livelist_condense, SPA

Incremented each time livelist condensing is canceled while in spa_livelist_condense_sync(). This option is used by the test suite to track race conditions.

zfs_livelist_condense_sync_pause

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

livelist, livelist_condense, SPA

When set, the livelist condense process pauses indefinitely before executing the synctask — spa_livelist_condense_sync(). This option is used by the test suite to trigger race conditions.

zfs_livelist_condense_zthr_cancel

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Change:

Dynamic

Tags:

livelist, livelist_condense, SPA

Incremented each time livelist condensing is canceled while in spa_livelist_condense_cb(). This option is used by the test suite to track race conditions.

zfs_livelist_condense_zthr_pause

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

livelist, livelist_condense, SPA

When set, the livelist condense process pauses indefinitely before executing the open context condensing work in spa_livelist_condense_cb(). This option is used by the test suite to trigger race conditions.

zfs_livelist_max_entries

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

500000

Change:

Dynamic

Tags:

DSL, dsl_deadlist, livelist

The threshold size (in block pointers) at which we create a new sub-livelist. Larger sublists are more costly from a memory perspective but the fewer sublists there are, the lower the cost of insertion.

zfs_livelist_min_percent_shared

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

75

Change:

Dynamic

Tags:

DSL, dsl_deadlist, livelist

If the amount of shared space between a snapshot and its clone drops below this threshold, the clone turns off the livelist and reverts to the old deletion method. This is in place because livelists no long give us a benefit once a clone has been overwritten enough.

zfs_log_sm_blksz

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Default:

131072

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

This is used as the block size for the space maps used for the log space map feature. These space maps benefit from a bigger block size as we expect to be writing a lot of data to them at once.

zfs_lua_max_instrlimit

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

100000000

Range:

0 to MAX_ULONG

Change:

Dynamic

Tags:

channel_programs, lua, zcp

The maximum execution time limit that can be set for a ZFS channel program, specified as a number of Lua instructions.

When to change: to enforce a CPU usage limit on ZFS channel programs

zfs_lua_max_memlimit

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

104857600

Range:

0 to MAX_ULONG

Change:

Dynamic

Tags:

channel_programs, lua, zcp

The maximum memory limit that can be set for a ZFS channel program, specified in bytes.

zfs_max_async_dedup_frees

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.4 - master

250000

v2.0 - v2.3

100,000

Change:

Dynamic

Tags:

dedup, DSL, dsl_scan

Maximum number of dedup, clone or gang blocks freed in a single TXG. These frees may require additional I/O, making them more expensive.

zfs_max_dataset_nesting

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

50

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

dataset

The maximum depth of nested datasets. This value can be tuned temporarily to fix existing datasets that exceed the predefined limit.

When to change: can be tuned temporarily to fix existing datasets that exceed the predefined limit

Notes: zfs_max_dataset_nesting limits the depth of nested datasets. Deeply nested datasets can overflow the stack. The maximum stack depth depends on kernel compilation options, so it is impractical to predict the possible limits. For kernels compiled with small stack sizes, zfs_max_dataset_nesting may require changes.

zfs_max_log_walking

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

5

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

The number of past TXGs that the flushing algorithm of the log spacemap feature uses to estimate incoming log blocks.

zfs_max_logsm_summary_length

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

10

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

Maximum number of rows allowed in the summary of the spacemap log.

zfs_max_missing_tvds

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

0

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

import, SPA

Number of missing top-level vdevs which will be allowed during pool import (only in read-only mode).

When to change: troubleshooting pools with missing devices

Notes: When importing a pool in readonly mode (zpool import -o readonly=on ...) then up to zfs_max_missing_tvds top-level vdevs can be missing, but the import can attempt to progress. Note: This is strictly intended for advanced pool recovery cases since missing data is almost inevitable. Pools with missing devices can only be imported read-only for safety reasons, and the pool’s failmode property is automatically set to continue The expected use case is to recover pool data immediately after accidentally adding a non-protected vdev to a protected pool. - With 1 missing top-level vdev, ZFS should be able to import the pool and mount all datasets. User data that was not modified after the missing device has been added should be recoverable. Thus snapshots created prior to the addition of that device should be completely intact. - With 2 missing top-level vdevs, some datasets may fail to mount since there are dataset statistics that are stored as regular metadata. Some data might be recoverable if those vdevs were added recently. - With 3 or more top-level missing vdevs, the pool is severely damaged and MOS entries may be missing entirely. Chances of data recovery are very low. Note that there are also risks of performing an inadvertent rewind as we might be missing all the vdevs with the latest uberblocks.

zfs_max_missing_tvds_cachefile

Versions:

master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

2

Change:

Dynamic

Tags:

SPA

Number of missing top-level vdevs tolerated when importing a pool from a cachefile, before the trusted config is read from the MOS. A cachefile can fall out of sync with the on-disk config after a device removal that did not rewrite the cachefile, so the default of 2 still lets the import reach a copy of the MOS.

zfs_max_missing_tvds_scan

Versions:

master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

0

Change:

Dynamic

Tags:

SPA

Number of missing top-level vdevs tolerated when importing a pool by scanning device paths, before the trusted config is read from the MOS. Defaults to 0 because a scan should detect every present device.

zfs_max_nvlist_src_size

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Change:

Dynamic

Tags:

ioctl

Maximum size in bytes allowed to be passed as zc_nvlist_src_size for ioctls on /dev/zfs. This prevents a user from causing the kernel to allocate an excessive amount of memory. When the limit is exceeded, the ioctl fails with EINVAL and a description of the error is sent to the zfs-dbgmsg log. This parameter should not need to be touched under normal circumstances. If 0, equivalent to a quarter of the user-wired memory limit under FreeBSD and to 134217728B (128 MiB) under Linux.

zfs_max_recordsize

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.2 - master

16777216

v0.6 - v2.1

1,048,576

Range:

512 to 16,777,216 (valid block sizes)

Change:

Dynamic

Tags:

DSL, dsl_dataset, filesystem, memory, volume

We currently support block sizes from 512 (512 B) to 16777216 (16 MiB). The benefits of larger blocks, and thus larger I/O, need to be weighed against the cost of COWing a giant block to modify one byte. Additionally, very large blocks can have an impact on I/O latency, and also potentially on the memory allocator. Therefore, we formerly forbade creating blocks larger than 1M. Larger blocks could be created by changing it, and pools with larger blocks can always be imported and used, regardless of this setting.

Note that it is still limited by default to 1 MiB on x86_32, because Linux’s 3/1 memory split doesn’t leave much room for 16M chunks.

When to change: To create datasets with larger volblocksize or recordsize

Notes: ZFS supports logical record (block) sizes from 512 bytes to 16 MiB. The benefits of larger blocks, and thus larger average I/O sizes, can be weighed against the cost of copy-on-write of large block to modify one byte. Additionally, very large blocks can have a negative impact on both I/O latency at the device level and the memory allocator. The zfs_max_recordsize parameter limits the upper bound of the dataset volblocksize and recordsize properties. Larger blocks can be created by enabling zpool large_blocks feature and changing this zfs_max_recordsize. Pools with larger blocks can always be imported and used, regardless of the value of zfs_max_recordsize. For 32-bit systems, zfs_max_recordsize also limits the size of kernel virtual memory caches used in the ZFS I/O pipeline (zio_buf_* and zio_data_buf_*). See also the zpool large_blocks feature.

zfs_mdcomp_disable

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Range:

0

compress metadata

1

do not compress metadata

Change:

Dynamic

Tags:

CPU, DMU, metadata

Disable meta data compression

Use 1 for yes and 0 for no (default).

When to change: When CPU cycles cost less than I/O

zfs_metaslab_condense_pct

Versions:

master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

200

Change:

Dynamic

Tags:

metaslab

Condense an on-disk space map when its size exceeds this percentage of the in-memory representation. The minimum is 100.

zfs_metaslab_find_max_tries

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

100

Change:

Dynamic

Tags:

metaslab

When not trying hard, we only consider this number of the best metaslabs. This improves performance, especially when there are many metaslabs per vdev and the allocation can’t actually be satisfied (so we would otherwise iterate all metaslabs).

zfs_metaslab_fragmentation_threshold

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.3 - master

77

v0.6 - v2.2

70

Range:

1 to 100

Change:

Dynamic

Tags:

allocation, fragmentation, metaslab, vdev

Allow metaslabs to keep their active state as long as their fragmentation percentage is no more than this value. An active metaslab that exceeds this threshold will no longer keep its active status allowing better metaslabs to be selected.

When to change: Testing metaslab allocation

Notes: Allow metaslabs to keep their active state as long as their fragmentation percentage is less than or equal to this value. When writing, an active metaslab whose fragmentation percentage exceeds zfs_metaslab_fragmentation_threshold is avoided allowing metaslabs with less fragmentation to be preferred. Metaslab fragmentation is used to calculate the overall pool fragmentation property value. However, individual metaslab fragmentation levels are observable using the zdb with the -mm option. zfs_metaslab_fragmentation_threshold works at the metaslab level and each top-level vdev has approximately zfs_vdev_default_ms_count metaslabs. See also zfs_mg_fragmentation_threshold

zfs_metaslab_max_size_cache_sec

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

3600

Change:

Dynamic

Tags:

metaslab

When we unload a metaslab, we cache the size of the largest free chunk. We use that cached size to determine whether or not to load a metaslab for a given allocation. As more frees accumulate in that metaslab while it’s unloaded, the cached max size becomes less and less accurate. After a number of seconds controlled by this tunable, we stop considering the cached max size and start considering only the histogram instead.

zfs_metaslab_mem_limit

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

25

Change:

Dynamic

Tags:

metaslab

When we are loading a new metaslab, we check the amount of memory being used to store metaslab range trees. If it is over a threshold, we attempt to unload the least recently used metaslab to prevent the system from clogging all of its memory with range trees. This tunable sets the percentage of total system memory that is the threshold.

zfs_metaslab_segment_weight_enabled

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not consider metaslab fragmentation

1

avoid metaslabs where free space is highly fragmented

Change:

Dynamic

Tags:

allocation, metaslab

Enable/disable segment-based metaslab selection.

When to change: When testing allocation and fragmentation

Notes: Enables metaslab allocation based on largest free segment rather than total amount of free space. The goal is to avoid metaslabs that exhibit free space fragmentation: when there is a lot of small free spaces, but few larger free spaces. If zfs_metaslab_segment_weight_enabled is enabled, then metaslab_fragmentation_factor_enabled is ignored.

zfs_metaslab_sm_blksz_no_log

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Default:

16384

Change:

Dynamic

Tags:

metaslab

Block size for the metaslab space maps in pools where the log_spacemap feature is disabled. Multiple metaslabs are modified per transaction group, so a smaller block size lets more, scattered I/O operations be issued. Must be a power of 2 greater than 4096. This parameter can only be set at module load time.

zfs_metaslab_sm_blksz_with_log

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Default:

131072

Change:

Dynamic

Tags:

metaslab

Block size for the metaslab space maps in pools where the log_spacemap feature is enabled. Changes are batched in the per-pool log spacemap and flushed to each metaslab’s space map only occasionally, so a larger block size is more efficient. Must be a power of 2 greater than 4096. This parameter can only be set at module load time.

zfs_metaslab_switch_threshold

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

2

Range:

0 to UINT64_MAX

Change:

Dynamic

Tags:

allocation, metaslab

When using segment-based metaslab selection, continue allocating from the active metaslab until this option’s worth of buckets have been exhausted.

When to change: When testing allocation and fragmentation

Notes: When using segment-based metaslab selection (see zfs_metaslab_segment_weight_enabled), continue allocating from the active metaslab until zfs_metaslab_switch_threshold worth of free space buckets have been exhausted.

zfs_metaslab_try_hard_before_gang

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

metaslab

  • If unset, we will first try normal allocation.

  • If that fails then we will do a gang allocation.

  • If that fails then we will do a “try hard” gang allocation.

  • If that fails then we will have a multi-layer gang block.

  • If set, we will first try normal allocation.

  • If that fails then we will do a “try hard” allocation.

  • If that fails we will do a gang allocation.

  • If that fails we will do a “try hard” gang allocation.

  • If that fails then we will have a multi-layer gang block.

zfs_mg_fragmentation_threshold

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v0.8 - master

95

v0.6 - v0.7

85

Range:

1 to 100

Change:

Dynamic

Tags:

allocation, fragmentation, metaslab, vdev

Metaslab groups are considered eligible for allocations if their fragmentation metric (measured as a percentage) is less than or equal to this value. If a metaslab group exceeds this threshold then it will be skipped unless all metaslab groups within the metaslab class have also crossed this threshold.

When to change: Testing metaslab allocation

Notes: Metaslab groups (top-level vdevs) are considered eligible for allocations if their fragmentation percentage metric is less than or equal to zfs_mg_fragmentation_threshold. If a metaslab group exceeds this threshold then it will be skipped unless all metaslab groups within the metaslab class have also crossed the zfs_mg_fragmentation_threshold threshold.

zfs_mg_noalloc_threshold

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0 to 100

Change:

Dynamic

Tags:

allocation, fragmentation, metaslab, vdev

Defines a threshold at which metaslab groups should be eligible for allocations. The value is expressed as a percentage of free space beyond which a metaslab group is always eligible for allocations. If a metaslab group’s free space is less than or equal to the threshold, the allocator will avoid allocating to that group unless all groups in the pool have reached the threshold. Once all groups have reached the threshold, all groups are allowed to accept allocations. The default value of 0 disables the feature and causes all metaslab groups to be eligible for allocations.

This parameter allows one to deal with pools having heavily imbalanced vdevs such as would be the case when a new vdev has been added. Setting the threshold to a non-zero percentage will stop allocations from being made to vdevs that aren’t filled to the specified percentage and allow lesser filled vdevs to acquire more allocations than they otherwise would under the old zfs_mg_alloc_failures facility.

When to change: To force rebalancing as top-level vdevs are added or expanded

Notes: Metaslab groups (top-level vdevs) with free space percentage greater than zfs_mg_noalloc_threshold are eligible for new allocations. If a metaslab group’s free space is less than or equal to the threshold, the allocator avoids allocating to that group unless all groups in the pool have reached the threshold. Once all metaslab groups have reached the threshold, all metaslab groups are allowed to accept allocations. The default value of 0 disables the feature and causes all metaslab groups to be eligible for allocations. This parameter allows one to deal with pools having heavily imbalanced vdevs such as would be the case when a new vdev has been added. Setting the threshold to a non-zero percentage will stop allocations from being made to vdevs that aren’t filled to the specified percentage and allow lesser filled vdevs to acquire more allocations than they otherwise would under the older zfs_mg_alloc_failures facility.

zfs_min_metaslabs_to_condense_pct

Versions:

master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

5

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

Minimum number of metaslabs to flush every TXG when condensing log spacemaps, as a percentage of the number of unflushed metaslabs at the time the zpool condense command was run.

zfs_min_metaslabs_to_flush

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

Minimum number of metaslabs to flush per dirty TXG.

zfs_multihost_fail_intervals

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v0.8 - master

10

v0.7

5

Change:

Dynamic

Tags:

import, MMP, multihost

Controls the behavior of the pool when multihost write failures or delays are detected.

When 0, multihost write failures or delays are ignored. The failures will still be reported to the ZED which depending on its configuration may take action such as suspending the pool or offlining a device.

Otherwise, the pool will be suspended if zfs_multihost_fail_intervals × zfs_multihost_interval milliseconds pass without a successful MMP write. This guarantees the activity test will see MMP writes if the pool is imported. 1 is equivalent to 2; this is necessary to prevent the pool from being suspended due to normal, small I/O latency variations.

Notes: zfs_multihost_fail_intervals controls the behavior of the pool when write failures are detected in the multihost multimodifier protection (MMP) subsystem. If zfs_multihost_fail_intervals = 0 then multihost write failures are ignored. The write failures are reported to the ZFS event daemon (zed) which can take action such as suspending the pool or offlining a device. If zfs_multihost_fail_intervals > 0 then sequential multihost write failures will cause the pool to be suspended. This occurs when (zfs_multihost_fail_intervals * zfs_multihost_interval) milliseconds have passed since the last successful multihost write. This guarantees the activity test will see multihost writes if the pool is attempted to be imported by another system.

zfs_multihost_history

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

import, MMP, multihost, SPA, spa_stats

Historical statistics for this many latest multihost updates will be available in /proc/spl/kstat/zfs/⟨pool⟩/multihost.

When to change: When testing multihost feature

Notes: The pool multihost multimodifier protection (MMP) subsystem can record historical updates in the /proc/spl/kstat/zfs/POOL_NAME/multihost file for debugging purposes. The number of lines of history is determined by zfs_multihost_history.

zfs_multihost_import_intervals

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v0.8 - master

20

v0.7

10

Range:

1 to UINT_MAX

Change:

Dynamic

Tags:

import, MMP, multihost

Used to control the duration of the activity test on import. Smaller values of zfs_multihost_import_intervals will reduce the import time but increase the risk of failing to detect an active pool. The total activity check time is never allowed to drop below one second.

On import the activity check waits a minimum amount of time determined by zfs_multihost_interval × zfs_multihost_import_intervals, or the same product computed on the host which last had the pool imported, whichever is greater. The activity check time may be further extended if the value of MMP delay found in the best uberblock indicates actual multihost updates happened at longer intervals than zfs_multihost_interval. A minimum of 100 ms is enforced.

0 is equivalent to 1.

Notes: zfs_multihost_import_intervals controls the duration of the activity test on pool import for the multihost multimodifier protection (MMP) subsystem. The activity test can be expected to take a minimum time of (zfs_multihost_import_intervals * zfs_multihost_interval * random(25%)) milliseconds. The random period of up to 25% improves simultaneous import detection. For example, if two hosts are rebooted at the same time and automatically attempt to import the pool, then is is highly probable that one host will win. Smaller values of zfs_multihost_import_intervals reduces the import time but increases the risk of failing to detect an active pool. The total activity check time is never allowed to drop below one second. Note: the multihost protection feature applies to storage devices that can be shared between multiple systems.

zfs_multihost_interval

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1000

Range:

100 to ULONG_MAX

Change:

Dynamic

Tags:

import, MMP, multihost, vdev

Used to control the frequency of multihost writes which are performed when the multihost pool property is on. This is one of the factors used to determine the length of the activity check during import.

The multihost write period is zfs_multihost_interval / leaf-vdevs. On average a multihost write will be issued for each leaf vdev every zfs_multihost_interval milliseconds. In practice, the observed period can vary with the I/O load and this observed value is the delay which is stored in the uberblock.

When to change: To optimize pool import time against possibility of simultaneous import by another system

Notes: zfs_multihost_interval controls the frequency of multihost writes performed by the pool multihost multimodifier protection (MMP) subsystem. The multihost write period is (zfs_multihost_interval / number of leaf-vdevs) milliseconds. Thus on average a multihost write will be issued for each leaf vdev every zfs_multihost_interval milliseconds. In practice, the observed period can vary with the I/O load and this observed value is the delay which is stored in the uberblock. On import the multihost activity check waits a minimum amount of time determined by (zfs_multihost_interval * zfs_multihost_import_intervals) with a lower bound of 1 second. The activity check time may be further extended if the value of mmp delay found in the best uberblock indicates actual multihost updates happened at longer intervals than zfs_multihost_interval Note: the multihost protection feature applies to storage devices that can be shared between multiple systems.

zfs_multilist_num_sublists

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.1 - master

0

v0.7 - v2.0

4 or the number of online CPUs, whichever is greater

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

ARC

To allow more fine-grained locking, each ARC state contains a series of lists for both data and metadata objects. Locking is performed at the level of these “sub-lists”. This parameters controls the number of sub-lists per ARC state, and also applies to other uses of the multilist data structure.

If 0, equivalent to the greater of the number of online CPUs and 4.

Notes: To allow more fine-grained locking, each ARC state contains a series of lists (sublists) for both data and metadata objects. Locking is performed at the sublist level. This parameters controls the number of sublists per ARC state, and also applies to other uses of the multilist data structure.

zfs_no_scrub_io

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

perform scrub I/O

1

do not perform scrub I/O

Change:

Dynamic

Tags:

DSL, dsl_scan, scrub

Set to disable scrub I/O. This results in scrubs not actually scrubbing data and simply doing a metadata crawl of the pool instead.

When to change: Testing scrub feature

Notes: When zfs_no_scrub_io = 1 scrubs do not actually scrub data and simply doing a metadata crawl of the pool instead.

zfs_no_scrub_prefetch

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

prefetch scrub I/Os

1

do not prefetch scrub I/Os

Change:

Dynamic

Tags:

DSL, dsl_scan, prefetch, scrub

Set to disable block prefetching for scrubs.

When to change: Testing scrub feature

Notes: When zfs_no_scrub_prefetch = 1, prefetch is disabled for scrub I/Os.

zfs_nocacheflush

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

send cache flush commands

1

do not send cache flush commands

Change:

Dynamic

Tags:

disks, vdev

Disable cache flush operations on disks when writing. Setting this will cause pool corruption on power loss if a volatile out-of-order write cache is enabled.

When to change: If the storage device has nonvolatile cache, then disabling cache flush can save the cost of occasional cache flush commands

Notes: ZFS uses barriers (volatile cache flush commands) to ensure data is committed to permanent media by devices. This ensures consistent on-media state for devices where caches are volatile (eg HDDs). For devices with nonvolatile caches, the cache flush operation can be a no-op. However, in some RAID arrays, cache flushes can cause the entire cache to be flushed to the backing devices. To ensure on-media consistency, keep cache flush enabled.

zfs_nopwrite_enabled

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

disable NOP-write feature

1

enable NOP-write feature

Change:

Dynamic

Tags:

checksum, debug, DMU

Allow no-operation writes. The occurrence of nopwrites will further depend on other pool properties (i.a. the checksumming and compression algorithms).

Notes: The NOP-write feature is enabled by default when a cryptographically-secure checksum algorithm is in use by the dataset. zfs_nopwrite_enabled allows the NOP-write feature to be completely disabled.

zfs_object_mutex_size

Versions:

v0.7 - master

Platforms:

Linux

Type:

uint

Default:

64

Range:

1 to UINT_MAX

Change:

Dynamic

Tags:

debug, znode

Size of the znode hashtable used for holds.

Due to the need to hold locks on objects that may not exist yet, kernel mutexes are not created per-object and instead a hashtable is used where collisions will result in objects waiting when there is not actually contention on the same object.

When to change: Testing znode mutex array deadlocks

Notes: zfs_object_mutex_size facilitates resizing the the per-dataset znode mutex array for testing deadlocks therein.

zfs_obsolete_min_time_ms

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

500

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

delete, DSL, dsl_scan, remove

Similar to zfs_free_min_time_ms, but for cleanup of old indirection records for removed vdevs.

Notes: zfs_obsolete_min_time_ms is similar to zfs_free_min_time_ms and used for cleanup of old indirection records for vdevs removed using the zpool remove command.

zfs_override_estimate_recordsize

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0=do not override, 1 to MAX_ULONG

Change:

Dynamic

Tags:

DMU, dmu_send, send

Setting this variable overrides the default logic for estimating block sizes when doing a zfs send. The default heuristic is that the average block size will be the current recordsize. Override this value if most data in your dataset is not of that size and you require accurate zfs send size estimates.

When to change: if most data in your dataset is not of the current recordsize and you require accurate zfs send size estimates

Notes: zfs_override_estimate_recordsize overrides the default logic for estimating block sizes when doing a zfs send. The default heuristic is that the average block size will be the current recordsize.

zfs_pd_bytes_max

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

52428800

Range:

0 to INT32_MAX

Change:

Dynamic

Tags:

DMU, dmu_traverse, prefetch, send

The number of bytes which should be prefetched during a pool traversal, like zfs send or other data crawling operations.

Notes: zfs_pd_bytes_max limits the number of bytes prefetched during a pool traversal (eg zfs send or other data crawling operations). These prefetches are referred to as “prescient prefetches” and are always 100% hit rate. The traversal operations do not use the default data or metadata prefetcher.

zfs_per_txg_dirty_frees_percent

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

30

Range:

0 to 100

Change:

Dynamic

Tags:

delete, DMU

Control percentage of dirtied indirect blocks from frees allowed into one TXG. After this threshold is crossed, additional frees will wait until the next TXG. 0 disables this throttle.

When to change: For zfs receive workloads, consider increasing or disabling. See ZFS I/O (ZIO) Scheduler

Notes: zfs_per_txg_dirty_frees_percent as a percentage of zfs_dirty_data_max controls the percentage of dirtied blocks from frees in one txg. After the threshold is crossed, additional dirty blocks from frees wait until the next txg. Thus, when deleting large files, filling consecutive txgs with deletes/frees, does not throttle other, perhaps more important, writes. A side effect of this throttle can impact zfs receive workloads that contain a large number of frees and the ignore_hole_birth optimization is disabled. The symptom is that the receive workload causes an increase in the frequency of txg commits. The frequency of txg commits is observable via the otime column of /proc/spl/kstat/zfs/POOLNAME/txgs. Since txg commits also flush data from volatile caches in HDDs to media, HDD performance can be negatively impacted. Also, since the frees do not consume much bandwidth over the pipe, the pipe can appear to stall. Thus the overall progress of receives is slower than expected. A value of zero will disable this throttle.

zfs_prefetch_disable

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

prefetch enabled

1

prefetch disabled

Change:

Dynamic

Tags:

DMU, dmu_zfetch, prefetch

Disable predictive prefetch. Note that it leaves “prescient” prefetch (for, e.g., zfs send) intact. Unlike predictive prefetch, prescient prefetch never issues I/O that ends up not being needed, so it can’t hurt performance.

When to change: In some case where the workload is completely random reads, overall performance can be better if prefetch is disabled

Verification: prefetch efficacy is observed by zarcstat and zarcsummary (arcstat and arc_summary before 2.4.0), and the relevant entries in /proc/spl/kstat/zfs/arcstats

Notes: zfs_prefetch_disable controls the predictive prefetcher. Note that it leaves “prescient” prefetch (eg prefetch for zfs send) intact (see zfs_pd_bytes_max)

zfs_qat_checksum_disable

Versions:

v0.8 - master

Platforms:

Linux

Type:

int

Default:

0

Range:

0

use QAT acceleration if available

1

do not use QAT acceleration

Change:

Dynamic

Tags:

checksum, QAT, qat_crypt

Disable QAT hardware acceleration for SHA256 checksums. May be unset after the ZFS modules have been loaded to initialize the QAT hardware as long as support is compiled in and the QAT driver is present.

When to change: Testing QAT functionality

Notes: zfs_qat_checksum_disable controls the Intel QuickAssist Technology (QAT) driver providing hardware acceleration for checksums. When the QAT hardware is present and qat driver available, the default behaviour is to enable QAT.

zfs_qat_compress_disable

Versions:

v0.8 - master

Platforms:

Linux

Type:

int

Default:

0

Range:

0

use QAT acceleration if available

1

do not use QAT acceleration

Change:

Dynamic

Tags:

compression, QAT, qat_compress

Disable QAT hardware acceleration for gzip compression. May be unset after the ZFS modules have been loaded to initialize the QAT hardware as long as support is compiled in and the QAT driver is present.

When to change: Testing QAT functionality

Notes: zfs_qat_compress_disable controls the Intel QuickAssist Technology (QAT) driver providing hardware acceleration for gzip compression. When the QAT hardware is present and qat driver available, the default behaviour is to enable QAT.

zfs_qat_disable

Versions:

v0.7

Platforms:

Linux, FreeBSD

Type:

int

Range:

0

use QAT acceleration if available

1

do not use QAT acceleration

Change:

Dynamic

Tags:

compression, QAT, qat_compress

This tunable disables qat hardware acceleration for gzip compression. It is available only if qat acceleration is compiled in and qat driver is present.

Use 1 for yes and 0 for no (default).

When to change: Testing QAT functionality

Notes: zfs_qat_disable controls the Intel QuickAssist Technology (QAT) driver providing hardware acceleration for gzip compression. When the QAT hardware is present and qat driver available, the default behaviour is to enable QAT.

zfs_qat_encrypt_disable

Versions:

v0.8 - master

Platforms:

Linux

Type:

int

Default:

0

Range:

0

use QAT acceleration if available

1

do not use QAT acceleration

Change:

Dynamic

Tags:

encryption, QAT, qat_crypt

Disable QAT hardware acceleration for AES-GCM encryption. May be unset after the ZFS modules have been loaded to initialize the QAT hardware as long as support is compiled in and the QAT driver is present.

When to change: Testing QAT functionality

Notes: zfs_qat_encrypt_disable controls the Intel QuickAssist Technology (QAT) driver providing hardware acceleration for encryption. When the QAT hardware is present and qat driver available, the default behaviour is to enable QAT.

zfs_read_chunk_size

Versions:

v0.6 - v0.8

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

1,048,576

Range:

512 to ULONG_MAX

Change:

Dynamic

Tags:

filesystem, vnops

Bytes to read per chunk

Default value: 1,048,576.

Notes: zfs_read_chunk_size is the limit for ZFS filesystem reads. If an application issues a read() larger than zfs_read_chunk_size, then the read() is divided into multiple operations no larger than zfs_read_chunk_size

zfs_read_history

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

debug, SPA, spa_stats

Historical statistics for this many latest reads will be available in /proc/spl/kstat/zfs/⟨pool⟩/reads.

When to change: To observe read operation details

Notes: Historical statistics for the last zfs_read_history reads are available in /proc/spl/kstat/zfs/POOL_NAME/reads

zfs_read_history_hits

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

do not include data for ARC hits

1

include ARC hit data

Change:

Dynamic

Tags:

debug, SPA, spa_stats

Include cache hits in read history

When to change: To observe read operation details with ARC hits

Notes: When zfs_read_history> 0, zfs_read_history_hits controls whether ARC hits are displayed in the read history file, /proc/spl/kstat/zfs/POOL_NAME/reads

zfs_rebuild_max_segment

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1048576

Change:

Dynamic

Tags:

vdev, vdev_rebuild

Maximum read segment size to issue when sequentially resilvering a top-level vdev.

zfs_rebuild_scrub_enabled

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

scrub, vdev, vdev_rebuild

Automatically start a pool scrub when the last active sequential resilver completes in order to verify the checksums of all blocks which have been resilvered. This is enabled by default and strongly recommended.

zfs_rebuild_vdev_limit

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

67108864

Change:

Dynamic

Tags:

vdev, vdev_rebuild

Maximum amount of I/O that can be concurrently issued for a sequential resilver per leaf device, given in bytes.

zfs_reconstruct_indirect_combinations_max

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

4096

Range:

0=do not limit attempts, 1 to MAX_INT = limit for attempts

Change:

Dynamic

Tags:

vdev, vdev_indirect, vdev_removal

If an indirect split block contains more than this many possible unique combinations when being reconstructed, consider it too computationally expensive to check them all. Instead, try at most this many randomly selected combinations each time the block is accessed. This allows all segment copies to participate fairly in the reconstruction when all combinations cannot be checked and prevents repeated use of one bad copy.

Notes: After device removal, if an indirect split block contains more than zfs_reconstruct_indirect_combinations_max many possible unique combinations when being reconstructed, it can be considered too computationally expensive to check them all. Instead, at most zfs_reconstruct_indirect_combinations_max randomly-selected combinations are attempted each time the block is accessed. This allows all segment copies to participate fairly in the reconstruction when all combinations cannot be checked and prevents repeated use of one bad copy.

zfs_recover

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

normal operation

1

attempt recovery zpool import

Change:

Dynamic

Tags:

import, SPA

Set to attempt to recover from fatal errors. This should only be used as a last resort, as it typically results in leaked space, or worse.

When to change: zfs_recover should only be used as a last resort, as it typically results in leaked space, or worse

Verification: check output of dmesg and other logs for details

Notes: zfs_recover can be set to true (1) to attempt to recover from otherwise-fatal errors, typically caused by on-disk corruption. When set, calls to zfs_panic_recover() will turn into warning messages rather than calling panic()

zfs_recv_best_effort_corrective

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Change:

Dynamic

Tags:

DMU, dmu_recv, receive

When this variable is set to non-zero a corrective receive:

  1. Does not enforce the restriction of source & destination snapshot GUIDs matching.

  2. If there is an error during healing, the healing receive is not terminated instead it moves on to the next record.

zfs_recv_defer_batch_size

Versions:

master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

33554432

Change:

Dynamic

Tags:

DMU, dmu_recv, receive

When a non-raw incremental stream reallocates an object at a different dnode size, zfs receive must wait for the old dnode’s free to be written out before the object number can be reused. Rather than forcing one transaction group sync per reallocated object, affected records are parked, up to this many bytes at a time, and applied together behind a single sync. Parked records hold their payloads in memory, so a receive may use up to this much memory in addition to zfs_recv_queue_length. While records are parked, the saved resume state is held back at the first parked record, so interrupting a receive can require resending up to this much of the stream. Setting this to 0 restores the previous behavior of one sync per reallocated object.

zfs_recv_queue_ff

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Change:

Dynamic

Tags:

DMU, dmu_recv, receive

The fill fraction of the zfs receive queue. The fill fraction controls the timing with which internal threads are woken up.

zfs_recv_queue_length

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

16777216

Range:

Must be at least twice the maximum recordsize or volblocksize in use

Change:

Dynamic

Tags:

DMU, dmu_recv, receive

The maximum number of bytes allowed in the zfs receive queue. This value must be at least twice the maximum block size in use.

When to change: When using the largest recordsize or volblocksize (16 MiB), increasing can improve receive efficiency

zfs_recv_write_batch_size

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1048576

Change:

Dynamic

Tags:

DMU, dmu_recv, receive

The maximum amount of data, in bytes, that zfs receive will write in one DMU transaction. This is the uncompressed size, even when receiving a compressed send stream. This setting will not reduce the write size below a single block. Capped at a maximum of 32 MiB.

zfs_removal_ignore_errors

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

during device removal: 0 = hard errors are not ignored, 1 = hard errors are ignored

Change:

Dynamic

Tags:

vdev, vdev_removal

Ignore hard I/O errors during device removal. When set, if a device encounters a hard I/O error during the removal process the removal will not be canceled. This can result in a normally recoverable block becoming permanently damaged and is hence not recommended. This should only be used as a last resort when the pool cannot be returned to a healthy state prior to removing the device.

When to change: See description for caveat

Notes: When removing a device, zfs_removal_ignore_errors controls the process for handling hard I/O errors. When set, if a device encounters a hard IO error during the removal process the removal will not be cancelled. This can result in a normally recoverable block becoming permanently damaged and is not recommended. This should only be used as a last resort when the pool cannot be returned to a healthy state prior to removing the device.

zfs_removal_suspend_progress

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0 = do not suspend during vdev removal

Change:

Dynamic

Tags:

vdev, vdev_removal

This is used by the test suite so that it can ensure that certain actions happen while in the middle of a removal.

When to change: do not change

Notes: zfs_removal_suspend_progress is used during automated testing of the ZFS code to incease test coverage.

zfs_remove_max_segment

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

16777216

Range:

maximum of the physical block size of all vdevs in the pool to 16,777,216 bytes (16 MiB)

Change:

Dynamic

Tags:

remove, vdev, vdev_removal

The largest contiguous segment that we will attempt to allocate when removing a device. If there is a performance problem with attempting to allocate large blocks, consider decreasing this. The default value is also the maximum.

When to change: after removing a top-level vdev, consider decreasing if there is a performance degradation when attempting to allocate large blocks

Notes: zfs_remove_max_segment sets the largest contiguous segment that ZFS attempts to allocate when removing a vdev. This can be no larger than 16MB. If there is a performance problem with attempting to allocate large blocks, consider decreasing this. The value is rounded up to a power-of-2.

zfs_resilver_defer_percent

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Change:

Dynamic

Tags:

DSL, dsl_scan

If the ongoing resilver progress is below this threshold, a new resilver will restart from scratch instead of being deferred after the current one finishes, even if the resilver_defer feature is enabled.

zfs_resilver_delay

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Default:

2

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

delay, DSL, dsl_scan, resilver, ZIO_scheduler

Number of ticks to delay prior to issuing a resilver I/O operation when a non-resilver or non-scrub I/O operation has occurred within the past zfs_scan_idle ticks.

Default value: 2.

When to change: increasing can reduce impact of resilver workload on dynamic workloads

Notes: zfs_resilver_delay sets a time-based delay for resilver I/Os. This delay is in addition to the ZIO scheduler’s treatment of scrub workloads. See also zfs_scan_idle

zfs_resilver_disable_defer

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

allow resilver_defer to postpone new resilver operations

1

immediately restart resilver when needed

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver

Ignore the resilver_defer feature, causing an operation that would start a resilver to immediately restart the one in progress.

When to change: if resilver postponement is not desired due to overall resilver time constraints

Notes: zfs_resilver_disable_defer disables the resilver_defer pool feature. The resilver_defer feature allows ZFS to postpone new resilvers if an existing resilver is in progress.

zfs_resilver_min_time_ms

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.4 - master

1500

v0.6 - v2.3

3,000

Range:

1 to zfs_txg_timeout converted to milliseconds

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver

Resilvers are processed by the sync thread. While resilvering, it will spend at least this much time working on a resilver between TXG flushes.

When to change: In some resilvering cases, increasing zfs_resilver_min_time_ms can result in faster completion

Notes: Resilvers are processed by the sync thread in syncing context. While resilvering, ZFS spends at least zfs_resilver_min_time_ms time working on a resilver between txg commits. The zfs_txg_timeout tunable sets a nominal timeout value for the txg commits. By default, this timeout is 5 seconds and the zfs_resilver_min_time_ms is 3 seconds. However, many variables contribute to changing the actual txg times. The measured txg interval is observed as the otime column (in nanoseconds) in the /proc/spl/kstat/zfs/POOL_NAME/txgs file. See also zfs_txg_timeout and zfs_scan_min_time_ms (removed after v0.7)

zfs_sb_uuid

Versions:

master

Platforms:

Linux

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

vfsops

Set the filesystem UUID of a filesystem, snapshot, or clone at mount time, from the pool GUID and the dataset GUID, as the guid property in zfsprops(7) describes. Set to 0 to leave the UUID null, as releases without this parameter did, for example when stale overlayfs file handles block a mount after an upgrade; see the Filesystem UUIDs and Overlayfs section of zpool-reguid(8). A change applies to the filesystems that are mounted after the change. For the filesystems that mount at boot, set the parameter before the first mount, on the kernel command line as zfs.zfs_sb_uuid=0 or in a modprobe.d file as options zfs zfs_sb_uuid=0. When the root filesystem is on ZFS, the initramfs must include that file. This parameter only applies on Linux.

zfs_scan_blkstats

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

DSL, dsl_scan

When enabled, ZFS counts blocks by type and indirection level during a scrub. The counts are kept in memory for debugging and are not exposed to userspace.

zfs_scan_checkpoint_intval

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

7200

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub

To preserve progress across reboots, the sequential scan algorithm periodically needs to stop metadata scanning and issue all the verification I/O to disk. The frequency of this flushing is determined by this tunable.

Notes: To preserve progress across reboots the sequential scan algorithm periodically needs to stop metadata scanning and issue all the verifications I/Os to disk every zfs_scan_checkpoint_intval seconds.

zfs_scan_fill_weight

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

3

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub

This tunable affects how scrub and resilver I/O segments are ordered. A higher number indicates that we care more about how filled in a segment is, while a lower number indicates we care more about the size of the extent without considering the gaps within a segment. This value is only tunable upon module insertion. Changing the value afterwards will have no effect on scrub or resilver performance.

When to change: Testing sequential scrub and resilver

Notes: This tunable affects how scrub and resilver I/O segments are ordered. A higher number indicates that we care more about how filled in a segment is, while a lower number indicates we care more about the size of the extent without considering the gaps within a segment.

zfs_scan_idle

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Default:

50

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub, ZIO_scheduler

Idle window in clock ticks. During a scrub or a resilver, if a non-scrub or non-resilver I/O operation has occurred during this window, the next scrub or resilver operation is delayed by, respectively zfs_scrub_delay or zfs_resilver_delay ticks.

Default value: 50.

When to change: as part of a resilver/scrub tuning effort

Notes: When a non-scan I/O has occurred in the past zfs_scan_idle clock ticks, then zfs_resilver_delay or zfs_scrub_delay are enabled.

zfs_scan_ignore_errors

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

do not ignore errors

1

ignore errors during pool scrub or resilver

Change:

Dynamic

Tags:

resilver, vdev

If set, remove the DTL (dirty time list) upon completion of a pool scan (scrub), even if there were unrepairable errors. Intended to be used during pool repair or recovery to stop resilvering when the pool is next imported.

When to change: See description above

Notes: zfs_scan_ignore_errors allows errors discovered during scrub or resilver to be ignored. This can be tuned as a workaround to remove the dirty time list (DTL) when completing a pool scan. It is intended to be used during pool repair or recovery to prevent resilvering when the pool is imported.

zfs_scan_issue_strategy

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0

fs will use strategy 1 during normal verification and strategy 2 while taking a checkpoint

1

data is verified as sequentially as possible, given the amount of memory reserved for scrubbing (see zfs_scan_mem_lim_fact). This can improve scrub performance if the pool’s data is heavily fragmented

2

the largest mostly-contiguous chunk of found data is verified first. By deferring scrubbing of small segments, we may later find adjacent data to coalesce and increase the segment size

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub

Determines the order that data will be verified while scrubbing or resilvering:

1

Data will be verified as sequentially as possible, given the amount of memory reserved for scrubbing (see zfs_scan_mem_lim_fact). This may improve scrub performance if the pool’s data is very fragmented.

2

The largest mostly-contiguous chunk of found data will be verified first. By deferring scrubbing of small segments, we may later find adjacent data to coalesce and increase the segment size.

0

Use strategy 1 during normal verification and strategy 2 while taking a checkpoint.

Notes: zfs_scan_issue_strategy controls the order of data verification while scrubbing or resilvering.

zfs_scan_legacy

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

use new method: scrubs and resilvers will gather metadata in memory before issuing sequential I/O

1

use legacy algorithm will be used where I/O is initiated as soon as it is discovered

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub

If unset, indicates that scrubs and resilvers will gather metadata in memory before issuing sequential I/O. Otherwise indicates that the legacy algorithm will be used, where I/O is initiated as soon as it is discovered. Unsetting will not affect scrubs or resilvers that are already in progress.

When to change: In some cases, the new scan mode can consumer more memory as it collects and sorts I/Os; using the legacy algorithm can be more memory efficient at the expense of HDD read efficiency

Notes: Setting zfs_scan_legacy = 1 enables the legacy scan and scrub behavior instead of the newer sequential behavior.

zfs_scan_max_ext_gap

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

2097152

Range:

512 to ULONG_MAX

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub

Sets the largest gap in bytes between scrub/resilver I/O operations that will still be considered sequential for sorting purposes. Changing this value will not affect scrubs or resilvers that are already in progress.

Notes: zfs_scan_max_ext_gap limits the largest gap in bytes between scrub and resilver I/Os that will still be considered sequential for sorting purposes.

zfs_scan_mem_lim_fact

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Change:

Dynamic

Tags:

DSL, dsl_scan, memory, resilver, scrub

Maximum fraction of RAM used for I/O sorting by sequential scan algorithm. This tunable determines the hard limit for I/O sorting memory usage. When the hard limit is reached we stop scanning metadata and start issuing data verification I/O. This is done until we get below the soft limit.

Notes: zfs_scan_mem_lim_fact limits the maximum fraction of RAM used for I/O sorting by sequential scan algorithm. When the limit is reached scanning metadata is stopped and data verification I/O is started. Data verification I/O continues until the memory used by the sorting algorithm drops by zfs_scan_mem_lim_soft_fact Memory used by the sequential scan algorithm can be observed as the kmem sio_cache. This is visible from procfs as grep sio_cache /proc/slabinfo and can be monitored using slab-monitoring tools such as slabtop

zfs_scan_mem_lim_soft_fact

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub

The fraction of the hard limit used to determined the soft limit for I/O sorting by the sequential scan algorithm. When we cross this limit from below no action is taken. When we cross this limit from above it is because we are issuing verification I/O. In this case (unless the metadata scan is done) we stop issuing verification I/O and start scanning metadata again until we get to the hard limit.

Notes: zfs_scan_mem_lim_soft_fact sets the fraction of the hard limit, zfs_scan_mem_lim_fact, used to determined the RAM soft limit for I/O sorting by the sequential scan algorithm. After zfs_scan_mem_lim_fact has been reached, metadata scanning is stopped until the RAM usage drops by zfs_scan_mem_lim_soft_fact

zfs_scan_min_time_ms

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Default:

1,000

Range:

1 to zfs_txg_timeout converted to milliseconds

Change:

Dynamic

Tags:

DSL, dsl_scan, scrub

Scrubs are processed by the sync thread. While scrubbing it will spend at least this much time working on a scrub between txg flushes.

Default value: 1,000.

When to change: In some scrub cases, increasing zfs_scan_min_time_ms can result in faster completion

Notes: Scrubs are processed by the sync thread in syncing context. While scrubbing, ZFS spends at least zfs_scan_min_time_ms time working on a scrub between txg commits. See also zfs_txg_timeout and zfs_resilver_min_time_ms

zfs_scan_report_txgs

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

DSL, dsl_scan

When reporting resilver throughput and estimated completion time use the performance observed over roughly the last zfs_scan_report_txgs TXGs. When set to zero performance is calculated over the time between checkpoints.

zfs_scan_strict_mem_lim

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

normal scan behaviour

1

check hard memory limit strictly during scan

Change:

Dynamic

Tags:

DSL, dsl_scan, memory, resilver, scrub

Enforce tight memory limits on pool scans when a sequential scan is in progress. When disabled, the memory limit may be exceeded by fast disks.

When to change: Do not change

Notes: When scrubbing or resilvering, by default, ZFS checks to ensure it is not over the hard memory limit before each txg commit. If finer-grained control of this is needed zfs_scan_strict_mem_lim can be set to 1 to enable checking before scanning each block.

zfs_scan_suspend_progress

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

do not freeze scans

1

freeze scans

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub

Freezes a scrub/resilver in progress without actually pausing it. Intended for testing/debugging.

When to change: testing or debugging scan code

Notes: zfs_scan_suspend_progress causes a scrub or resilver scan to freeze without actually pausing.

zfs_scan_vdev_limit

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.1 - master

16777216

v0.8 - v2.0

41943040

Range:

512 to ULONG_MAX

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub, vdev

Maximum amount of data that can be concurrently issued at once for scrubs and resilvers per leaf device, given in bytes.

Notes: zfs_scan_vdev_limit is the maximum amount of data that can be concurrently issued at once for scrubs and resilvers per leaf vdev. zfs_scan_vdev_limit attempts to strike a balance between keeping the leaf vdev queues full of I/Os while not overflowing the queues causing high latency resulting in long txg sync times. While zfs_scan_vdev_limit represents a bandwidth limit, the existing I/O limit of zfs_vdev_scrub_max_active remains in effect, too.

zfs_scrub_after_expand

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

raidz, scrub, vdev

Automatically start a pool scrub after a RAIDZ expansion completes in order to verify the checksums of all blocks which have been copied during the expansion. This is enabled by default and strongly recommended.

zfs_scrub_delay

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Default:

4

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

delay, DSL, dsl_scan, scrub, ZIO_scheduler

Number of ticks to delay prior to issuing a scrub I/O operation when a non-scrub or non-resilver I/O operation has occurred within the past zfs_scan_idle ticks.

Default value: 4.

When to change: increasing can reduce impact of scrub workload on dynamic workloads

Notes: zfs_scrub_delay sets a time-based delay for scrub I/Os. This delay is in addition to the ZIO scheduler’s treatment of scrub workloads. See also zfs_scan_idle

zfs_scrub_error_blocks_per_txg

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

4096

Change:

Dynamic

Tags:

DSL, dsl_scan, scrub

Error blocks to be scrubbed in one txg.

zfs_scrub_min_time_ms

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.4 - master

750

v0.8 - v2.3

1,000

Range:

1 to (zfs_txg_timeout - 1)

Change:

Dynamic

Tags:

DSL, dsl_scan, scrub

Scrubs are processed by the sync thread. While scrubbing, it will spend at least this much time working on a scrub between TXG flushes.

Notes: Scrubs are processed by the sync thread. While scrubbing at least zfs_scrub_min_time_ms time is spent working on a scrub between txg syncs.

zfs_scrub_partial_writes

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

raidz, scrub, vdev

If a write to a multi-disk vdev fails, but the data is recoverable, the data is persisted on disk but may not be as redundant as the vdev usually ensures. If this tunable is set, we issue a read after such a write error to detect the full extent of the problem and attempt to recover from it. Note: This currently only works with RAID-Z and dRAID.

zfs_send_corrupt_data

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

do not send corrupt data

1

replace corrupt data with cookie

Change:

Dynamic

Tags:

DMU, dmu_send, send

Allow sending of corrupt data (ignore read/checksum errors when sending).

When to change: When data corruption exists and an attempt to recover at least some data via zfs send is needed

Notes: zfs_send_corrupt_data enables zfs send to send of corrupt data by ignoring read and checksum errors. The corrupted or unreadable blocks are replaced with the value 0x2f5baddb10c (ZFS bad block)

zfs_send_no_prefetch_queue_ff

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Change:

Dynamic

Tags:

DMU, dmu_send, prefetch, send

The fill fraction of the zfs send internal queues. The fill fraction controls the timing with which internal threads are woken up.

zfs_send_no_prefetch_queue_length

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1048576

Change:

Dynamic

Tags:

DMU, dmu_send, prefetch, send

The maximum number of bytes allowed in zfs send’s internal queues.

zfs_send_queue_ff

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Change:

Dynamic

Tags:

DMU, dmu_send, send

The fill fraction of the zfs send prefetch queue. The fill fraction controls the timing with which internal threads are woken up.

zfs_send_queue_length

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

16777216

Range:

Must be at least twice the maximum recordsize or volblocksize in use

Change:

Dynamic

Tags:

DMU, dmu_send, send

The maximum number of bytes allowed that will be prefetched by zfs send. This value must be at least twice the maximum block size in use.

When to change: When using the largest recordsize or volblocksize (16 MiB), increasing can improve send efficiency

Notes: zfs_send_queue_length is the maximum number of bytes allowed in the zfs send queue.

zfs_send_unmodified_spill_blocks

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not send unmodified spill blocks

1

send unmodified spill blocks

Change:

Dynamic

Tags:

DMU, dmu_send, send

Include unmodified spill blocks in the send stream. Under certain circumstances, previous versions of ZFS could incorrectly remove the spill block from an existing object. Including unmodified copies of the spill blocks creates a backwards-compatible stream which will recreate a spill block if it was incorrectly removed.

Notes: zfs_send_unmodified_spill_blocks enables sending of unmodified spill blocks in the send stream. Under certain circumstances, previous versions of ZFS could incorrectly remove the spill block from an existing object. Including unmodified copies of the spill blocks creates a backwards compatible stream which will recreate a spill block if it was incorrectly removed.

zfs_slow_io_events_per_second

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

20

Range:

zed threshold to MAX_UINT

Change:

Dynamic

Tags:

vdev

Rate limit delay zevents (which report slow I/O operations) to this many per second.

Notes: zfs_slow_io_events_per_second is a rate limit for slow I/O events. Note that this should not be set below the zed thresholds (currently 10 checksums over 10 sec) or else zed may not trigger any action.

zfs_snapshot_history_enabled

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

DSL, dsl_dataset, snapshot

Whether snapshot creation and destruction events are recorded in the pool history log, viewable with zpool history.

zfs_snapshot_no_setuid

Versions:

v2.3 - master

Platforms:

Linux

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

ctldir, snapshot

Whether to disable setuid/setgid support for snapshot mounts triggered by access to the .zfs/snapshot directory by setting the nosuid mount option.

zfs_spa_discard_memory_limit

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

16777216

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

checkpoint, SPA

Maximum memory used for prefetching a checkpoint’s space map on each vdev while discarding the checkpoint.

zfs_special_class_metadata_reserve_pct

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

25

Range:

0 to 100

Change:

Dynamic

Tags:

metadata, SPA, special_vdev

Only allow small data blocks to be allocated on the special and dedup vdev types when the available free space percentage on these vdevs exceeds this value. This ensures reserved space is available for pool metadata as the special vdevs approach capacity.

Notes: zfs_special_class_metadata_reserve_pct sets a threshold for space in special vdevs to be reserved exclusively for metadata. This prevents small data blocks from completely consuming a special vdev.

zfs_sync_pass_deferred_free

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

2

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

SPA, ZIO

Flushing of data to disk is done in passes. Defer frees starting in this pass.

When to change: Testing SPA sync process

Notes: The SPA sync process is performed in multiple passes. Once the pass number reaches zfs_sync_pass_deferred_free, frees are no long processed and must wait for the next SPA sync. The zfs_sync_pass_deferred_free value is expected to be removed as a tunable once the optimal value is determined during field testing. The zfs_sync_pass_deferred_free pass must be greater than 1 to ensure that regular blocks are not deferred.

zfs_sync_pass_dont_compress

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v0.8 - master

8

v0.6 - v0.7

5

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

SPA, ZIO

Starting in this sync pass, disable compression (including of metadata). With the default setting, in practice, we don’t have this many sync passes, so this has no effect.

The original intent was that disabling compression would help the sync passes to converge. However, in practice, disabling compression increases the average number of sync passes; because when we turn compression off, many blocks’ size will change, and thus we have to re-allocate (not overwrite) them. It also increases the number of 128 KiB allocations (e.g. for indirect blocks and spacemaps) because these will not be compressed. The 128 KiB allocations are especially detrimental to performance on highly fragmented systems, which may have very few free segments of this size, and may need to load new metaslabs to satisfy these allocations.

When to change: Testing SPA sync process

Notes: The SPA sync process is performed in multiple passes. Once the pass number reaches zfs_sync_pass_dont_compress, data block compression is no longer processed and must wait for the next SPA sync. The zfs_sync_pass_dont_compress value is expected to be removed as a tunable once the optimal value is determined during field testing.

zfs_sync_pass_rewrite

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

2

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

SPA, ZIO

Rewrite new block pointers starting in this pass.

When to change: Testing SPA sync process

Notes: The SPA sync process is performed in multiple passes. Once the pass number reaches zfs_sync_pass_rewrite, blocks can be split into gang blocks. The zfs_sync_pass_rewrite value is expected to be removed as a tunable once the optimal value is determined during field testing.

zfs_sync_taskq_batch_pct

Versions:

v0.7 - v2.2

Platforms:

Linux, FreeBSD

Type:

int

Default:

75

Range:

1 to 100

Change:

Dynamic

Tags:

DSL, dsl_pool, SPA, taskq

This controls the number of threads used by dp_sync_taskq. The default value of 75% will create a maximum of one thread per CPU.

When to change: to adjust the number of dp_sync_taskq threads

Notes: zfs_sync_taskq_batch_pct controls the number of threads used by the DSL pool sync taskq, dp_sync_taskq

zfs_top_maxinflight

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Default:

32

Range:

1 to MAX_INT

Change:

Dynamic

Tags:

DSL, dsl_scan, resilver, scrub, ZIO_scheduler

Max concurrent I/Os per top-level vdev (mirrors or raidz arrays) allowed during scrub or resilver operations.

Default value: 32.

When to change: for modern ZFS versions, the ZIO scheduler limits usually take precedence

Notes: zfs_top_maxinflight is used to limit the maximum number of I/Os queued to top-level vdevs during scrub or resilver operations. The actual top-level vdev limit is calculated by multiplying the number of child vdevs by zfs_top_maxinflight This limit is an additional cap over and above the scan limits

zfs_traverse_indirect_prefetch_limit

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

32

Change:

Dynamic

Tags:

DMU, dmu_traverse, prefetch

The number of blocks pointed by indirect (non-L0) block which should be prefetched during a pool traversal, like zfs send or other data crawling operations.

zfs_trim_extent_bytes_max

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

134217728

Range:

zfs_trim_extent_bytes_min to MAX_UINT

Change:

Dynamic

Tags:

trim, vdev, vdev_trim

Maximum size of TRIM command. Larger ranges will be split into chunks no larger than this value before issuing.

When to change: if the device can efficiently handle larger trim requests

Notes: zfs_trim_extent_bytes_max sets the maximum size of a trim (aka discard, scsi unmap) command. Ranges larger than zfs_trim_extent_bytes_max are split in to chunks no larger than zfs_trim_extent_bytes_max bytes prior to being issued to the device. Use zpool iostat -w to observe the latency of trim commands.

zfs_trim_extent_bytes_min

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

32768

Range:

0=trim all unallocated space, otherwise minimum physical block size to UINT_MAX

Change:

Dynamic

Tags:

trim, vdev, vdev_trim

Minimum size of TRIM commands. TRIM ranges smaller than this will be skipped, unless they’re part of a larger range which was chunked. This is done because it’s common for these small TRIMs to negatively impact overall performance.

When to change: when trim is in use and device performance suffers from trimming small allocations

Notes: zfs_trim_extent_bytes_min sets the minimum size of trim (aka discard, scsi unmap) commands. Trim ranges smaller than zfs_trim_extent_bytes_min are skipped unless they’re part of a larger range which was broken in to chunks. Some devices have performance degradation during trim operations, so using a larger zfs_trim_extent_bytes_min can reduce the total amount of space trimmed. Use zpool iostat -w to observe the latency of trim commands.

zfs_trim_metaslab_skip

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0

do not skip uninitialized metaslabs during trim

1

skip uninitialized metaslabs during trim

Change:

Dynamic

Tags:

metaslab, trim, vdev, vdev_trim

Skip uninitialized metaslabs during the TRIM process. This option is useful for pools constructed from large thinly-provisioned devices where TRIM operations are slow. As a pool ages, an increasing fraction of the pool’s metaslabs will be initialized, progressively degrading the usefulness of this option. This setting is stored when starting a manual TRIM and will persist for the duration of the requested TRIM.

Notes: zfs_trim_metaslab_skip enables uninitialized metaslabs to be skipped during the trim (aka discard, scsi unmap) process. zfs_trim_metaslab_skip can be useful for pools constructed from large thinly-provisioned devices where trim operations perform slowly. As a pool ages an increasing fraction of the pool’s metaslabs are initialized, progressively degrading the usefulness of this option. This setting is stored when starting a manual trim and persists for the duration of the requested trim. Use zpool iostat -w to observe the latency of trim commands.

zfs_trim_queue_limit

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Range:

1 to MAX_UINT

Change:

Dynamic

Tags:

trim, vdev, vdev_trim

Maximum number of queued TRIMs outstanding per leaf vdev. The number of concurrent TRIM commands issued to the device is controlled by zfs_vdev_trim_min_active and zfs_vdev_trim_max_active.

When to change: to restrict the number of trim commands in the queue

Notes: zfs_trim_queue_limit sets the maximum queue depth for leaf vdevs. See also zfs_vdev_trim_max_active and zfs_trim_extent_bytes_max Use zpool iostat -q to observe trim queue depth.

zfs_trim_txg_batch

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

32

Range:

1 to MAX_UINT

Change:

Dynamic

Tags:

trim, vdev, vdev_trim

The number of transaction groups’ worth of frees which should be aggregated before TRIM operations are issued to the device. This setting represents a trade-off between issuing larger, more efficient TRIM operations and the delay before the recently trimmed space is available for use by the device.

Increasing this value will allow frees to be aggregated for a longer time. This will result is larger TRIM operations and potentially increased memory usage. Decreasing this value will have the opposite effect. The default of 32 was determined to be a reasonable compromise.

Notes: zfs_trim_txg_batch sets the number of transaction groups worth of frees which should be aggregated before trim (aka discard, scsi unmap) commands are issued to a device. This setting represents a trade-off between issuing larger, more efficient trim commands and the delay before the recently trimmed space is available for use by the device. Increasing this value will allow frees to be aggregated for a longer time. This will result is larger trim operations and potentially increased memory usage. Decreasing this value will have the opposite effect. The default value of 32 was empirically determined to be a reasonable compromise.

zfs_txg_history

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.3 - master

100

v0.6 - v2.2

0

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

debug, SPA, spa_stats, TXG

Historical statistics for this many latest TXGs will be available in /proc/spl/kstat/zfs/⟨pool⟩/TXGs.

When to change: To observe details of SPA sync behavior.

Notes: Historical statistics for the last zfs_txg_history txg commits are available in /proc/spl/kstat/zfs/POOL_NAME/txgs The work required to measure the txg commit (SPA statistics) is low. However, for debugging purposes, it can be useful to observe the SPA statistics.

zfs_txg_timeout

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

5

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

SPA, TXG, ZIO_scheduler

Flush dirty data to disk at least every this many seconds (maximum TXG duration).

When to change: To optimize the work done by txg commit relative to the pool requirements. See also section ZFS I/O Scheduler

Notes: The open txg is committed to the pool periodically (SPA sync) and zfs_txg_timeout represents the default target upper limit. txg commits can occur more frequently and a rapid rate of txg commits often indicates a busy write workload, quota limits reached, or the free space is critically low. Many variables contribute to changing the actual txg times. txg commits can also take longer than zfs_txg_timeout if the ZFS write throttle is not properly tuned or the time to sync is otherwise delayed (eg slow device). Shorter txg commit intervals can occur due to zfs_dirty_data_sync_percent for write-intensive workloads. The measured txg interval is observed as the otime column (in nanoseconds) in the /proc/spl/kstat/zfs/POOL_NAME/txgs file. See also zfs_dirty_data_sync_percent and zfs_txg_history

zfs_unflushed_log_block_max

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.1 - master

131072

v2.0

262144 (256K)

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

Describes the maximum number of log spacemap blocks allowed for each pool. The default value means that the space in all the log spacemaps can add up to no more than 131072 blocks (which means 16 GiB of logical space before compression and ditto blocks, assuming that blocksize is 128 KiB).

This tunable is important because it involves a trade-off between import time after an unclean export and the frequency of flushing metaslabs. The higher this number is, the more log blocks we allow when the pool is active which means that we flush metaslabs less often and thus decrease the number of I/O operations for spacemap updates per TXG. At the same time though, that means that in the event of an unclean export, there will be more log spacemap blocks for us to read, inducing overhead in the import time of the pool. The lower the number, the amount of flushing increases, destroying log blocks quicker as they become obsolete faster, which leaves less blocks to be read during import time after a crash.

Each log spacemap block existing during pool import leads to approximately one extra logical I/O issued. This is the reason why this tunable is exposed in terms of blocks rather than space used.

zfs_unflushed_log_block_min

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1000

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

If the number of metaslabs is small and our incoming rate is high, we could get into a situation that we are flushing all our metaslabs every TXG. Thus we always allow at least this many log blocks.

zfs_unflushed_log_block_pct

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

400

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

Tunable used to determine the number of blocks that can be used for the spacemap log, expressed as a percentage of the total number of unflushed metaslabs in the pool.

zfs_unflushed_log_txg_max

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1000

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

Tunable limiting maximum time in TXGs any metaslab may remain unflushed. It effectively limits maximum number of unflushed per-TXG spacemap logs that need to be read after unclean pool export.

zfs_unflushed_max_mem_amt

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1073741824

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

Upper-bound limit for unflushed metadata changes to be held by the log spacemap in memory, in bytes.

zfs_unflushed_max_mem_ppm

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

1000

Change:

Dynamic

Tags:

SPA, spa_log_spacemap

Part of overall system memory that ZFS allows to be used for unflushed metadata changes by the log spacemap, in millionths.

zfs_user_indirect_is_special

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not write user indirect blocks to a special vdev

1

write user indirect blocks to a special vdev

Change:

Dynamic

Tags:

SPA, special_vdev

If enabled, ZFS will place user data indirect blocks into the special allocation class.

When to change: to force user data indirect blocks to remain in the main pool top-level vdevs

Notes: If special vdevs are in use, zfs_user_indirect_is_special enables user data indirect blocks (a form of metadata) to be written to the special vdevs.

zfs_vdev_aggregate_trim

Versions:

v0.8 - v2.1

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

do not attempt to aggregate trim commands

1

attempt to aggregate trim commands

Change:

Dynamic

Tags:

trim, vdev, vdev_queue, ZIO_scheduler

Allow TRIM I/Os to be aggregated. This is normally not helpful because the extents to be trimmed will have been already been aggregated by the metaslab. This option is provided for debugging and performance analysis.

When to change: when debugging trim code or trim performance issues

Notes: zfs_vdev_aggregate_trim allows trim I/Os to be aggregated. This is normally not helpful because the extents to be trimmed will have been already been aggregated by the metaslab.

zfs_vdev_aggregation_limit

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v0.8 - master

1,048,576

v0.6 - v0.7

131,072

Range:

0 to 1,048,576 (default) or 16,777,216 (if zpool large_blocks feature is enabled)

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Max vdev I/O aggregation size.

When to change: If the workload does not benefit from aggregation, the zfs_vdev_aggregation_limit can be reduced to avoid aggregation attempts

Verification: ZFS aggregation is observed with zpool iostat -r and the block scheduler merging is observed with iostat -x

Notes: To reduce IOPs, small, adjacent I/Os can be aggregated (coalesced) into a large I/O. For reads, aggregations occur across small adjacency gaps. For writes, aggregation can occur at the ZFS or disk level. zfs_vdev_aggregation_limit is the upper bound on the size of the larger, aggregated I/O. Setting zfs_vdev_aggregation_limit = 0 effectively disables aggregation by ZFS. However, the block device scheduler can still merge (aggregate) I/Os. Also, many devices, such as modern HDDs, contain schedulers that can aggregate I/Os. In general, I/O aggregation can improve performance for devices, such as HDDs, where ordering I/O operations for contiguous LBAs is a benefit. For random access devices, such as SSDs, aggregation might not improve performance relative to the CPU cycles needed to aggregate. For devices that represent themselves as having no rotation, the zfs_vdev_aggregation_limit_non_rotating parameter is used instead of zfs_vdev_aggregation_limit

zfs_vdev_aggregation_limit_non_rotating

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

131072

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Max vdev I/O aggregation size for non-rotating media.

When to change: see zfs_vdev_aggregation_limit

Notes: zfs_vdev_aggregation_limit_non_rotating is the equivalent of zfs_vdev_aggregation_limit for devices which represent themselves as non-rotating to the Linux blkdev interfaces. Such devices have a value of 0 in /sys/block/DEVICE/queue/rotational and are expected to be SSDs.

zfs_vdev_async_read_max_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

3

Range:

1 to zfs_vdev_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Maximum asynchronous read I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

zfs_vdev_async_read_min_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Range:

1 to ( zfs_vdev_async_read_max_active - 1)

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Minimum asynchronous read I/O operation active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

zfs_vdev_async_write_active_max_dirty_percent

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

60

Range:

0 to 100

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

When the pool has more than this much dirty data, use zfs_vdev_async_write_max_active to limit active async writes. If the dirty data is between the minimum and maximum, the active I/O limit is linearly interpolated. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Sch eduler

Notes: When the amount of dirty data exceeds the threshold zfs_vdev_async_write_active_max_dirty_percent of zfs_dirty_data_max dirty data, then zfs_vdev_async_write_max_active is used to limit active async writes. If the dirty data is between zfs_vdev_async_write_active_min_dirty_percent and zfs_vdev_async_write_active_max_dirty_percent, the active I/O limit is linearly interpolated between zfs_vdev_async_write_min_active and zfs_vdev_async_write_max_active

zfs_vdev_async_write_active_min_dirty_percent

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

30

Range:

0 to (zfs_vdev_async_write_active_max_dirty_percent - 1)

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

When the pool has less than this much dirty data, use zfs_vdev_async_write_min_active to limit active async writes. If the dirty data is between the minimum and maximum, the active I/O limit is linearly interpolated. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Sch eduler

Notes: If the amount of dirty data is between zfs_vdev_async_write_active_min_dirty_percent and zfs_vdev_async_write_active_max_dirty_percent of zfs_dirty_data_max, the active I/O limit is linearly interpolated between zfs_vdev_async_write_min_active and zfs_vdev_async_write_max_active

zfs_vdev_async_write_max_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.2 - master

10

v2.1

30

v0.6 - v2.0

10

Range:

1 to zfs_vdev_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Maximum asynchronous write I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O S cheduler

zfs_vdev_async_write_min_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v0.7 - master

2

v0.6

1

Range:

1 to zfs_vdev_async_write_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Minimum asynchronous write I/O operations active to each device. See ZFS I/O SCHEDULER.

Lower values are associated with better latency on rotational media but poorer resilver performance. The default value of 2 was chosen as a compromise. A value of 3 has been shown to improve resilver performance further at a cost of further increasing latency.

When to change: See ZFS I/O S cheduler

Notes: zfs_vdev_async_write_min_active sets the minimum asynchronous write I/Os active to each device. Lower values are associated with better latency on rotational media but poorer resilver performance. The default value of 2 was chosen as a compromise. A value of 3 has been shown to improve resilver performance further at a cost of further increasing latency.

zfs_vdev_cache_bshift

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

int

Default:

16

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

vdev, vdev_cache

Shift size to inflate reads to.

When to change: Do not change

Notes: Note: with the current ZFS code, the vdev cache is not helpful and in some cases actually harmful. Thus it is disabled by setting the zfs_vdev_cache_size to zero. This related tunable is, by default, inoperative. All read I/Os smaller than zfs_vdev_cache_max are turned into (1 << zfs_vdev_cache_bshift) byte reads by the vdev cache. At most zfs_vdev_cache_size bytes will be kept in each vdev’s cache.

zfs_vdev_cache_max

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

int

Default:

16384

Range:

512 to INT_MAX

Change:

Dynamic

Tags:

vdev, vdev_cache

Inflate reads smaller than this value to meet the zfs_vdev_cache_bshift size (default 64kB).

When to change: Do not change

Notes: Note: with the current ZFS code, the vdev cache is not helpful and in some cases actually harmful. Thus it is disabled by setting the zfs_vdev_cache_size to zero. This related tunable is, by default, inoperative. All read I/Os smaller than zfs_vdev_cache_max will be turned into (1 <<zfs_vdev_cache_bshift byte reads by the vdev cache. At most zfs_vdev_cache_size bytes will be kept in each vdev’s cache.

zfs_vdev_cache_size

Versions:

v0.6 - v2.1

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 to MAX_INT

Change:

Prior to module load

Tags:

vdev, vdev_cache

Total size of the per-disk cache in bytes.

Currently this feature is disabled, as it has been found to not be helpful for performance and in some cases harmful.

When to change: Do not change

Verification: vdev cache statistics are available in the /proc/spl/kstat/zfs/vdev_cache_stats file

Notes: Note: with the current ZFS code, the vdev cache is not helpful and in some cases actually harmful. Thusit is disabled by setting the zfs_vdev_cache_size = 0 zfs_vdev_cache_size is the size of the vdev cache.

zfs_vdev_def_queue_depth

Versions:

v2.2 - v2.3

Platforms:

Linux, FreeBSD

Type:

uint

Default:

32

Change:

Dynamic

Tags:

vdev, vdev_queue

Default queue depth for each vdev IO allocator. Higher values allow for better coalescing of sequential writes before sending them to the disk, but can increase transaction commit times.

zfs_vdev_default_ms_count

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

200

Range:

16 to MAX_INT

Change:

Dynamic

Tags:

allocation, vdev

When a vdev is added, target this number of metaslabs per top-level vdev.

When to change: for development testing purposes only

zfs_vdev_default_ms_shift

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

29

Change:

Dynamic

Tags:

vdev

Default lower limit for metaslab size.

zfs_vdev_direct_write_verify

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

Linux

Change:

Dynamic

Tags:

vdev

If non-zero, then a Direct I/O write’s checksum will be verified every time the write is issued and before it is committed to the block pointer. In the event the checksum is not valid then the I/O operation will return EIO. This module parameter can be used to detect if the contents of the users buffer have changed in the process of doing a Direct I/O write. It can also help to identify if reported checksum errors are tied to Direct I/O writes. Each verify error causes a dio_verify_wr zevent. Direct Write I/O checksum verify errors can be seen with zpool status -d. The default value for this is 1 on Linux, but is 0 for FreeBSD because user pages can be placed under write protection in FreeBSD before the Direct I/O write is issued.

zfs_vdev_disk_calling_thread_io

Versions:

master

Platforms:

Linux

Type:

uint

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

vdev, vdev_disk

Controls calling thread io, note that we only wait for the zio to complete if it bypassed the vdev queue, all this module parameter does is enable that capability. May lead to performance improvements when enabled if backing vdev devices are fast and low latency. May impact performance with certain workloads when enabled on raidz or draid zpool configurations. This parameter currently only applies on Linux.

zfs_vdev_disk_classic

Versions:

v2.2 - v2.3

Platforms:

Linux

Type:

uint

Default:

0

Range:

0 | 1

Change:

Prior to module load

Tags:

vdev, vdev_disk

If set to 1, OpenZFS will submit IO to Linux using the method it used in 2.2 and earlier. This “classic” method has known issues with highly fragmented IO requests and is slower on many workloads, but it has been in use for many years and is known to be very stable. If you set this parameter, please also open a bug report why you did so, including the workload involved and any error messages.

This parameter and the classic submission method will be removed once we have total confidence in the new method.

This parameter only applies on Linux, and can only be set at module load time.

zfs_vdev_disk_max_segs

Versions:

v2.2 - master

Platforms:

Linux

Type:

uint

Default:

0

Change:

Dynamic

Tags:

vdev, vdev_disk

Maximum number of segments to add to a BIO (min 4). If this is higher than the maximum allowed by the device queue or the kernel itself, it will be clamped. Setting it to zero will cause the kernel’s ideal size to be used. This parameter only applies on Linux.

zfs_vdev_dtl_sm_blksz

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Default:

4096

Change:

Dynamic

Tags:

vdev

Since the DTL space map of a vdev is not expected to have a lot of entries, we default its block size to 4K.

zfs_vdev_failfast_mask

Versions:

v2.2 - master

Platforms:

Linux

Type:

uint

Default:

1

Change:

Dynamic

Tags:

vdev, vdev_disk

Defines if the driver should retire on a given error type. The following options may be bitwise-ored together:

Value

Name

Description

1

Device

No driver retries on device errors

2

Transport

No driver retries on transport errors.

4

Driver

No driver retries on driver errors.

zfs_vdev_initializing_max_active

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Range:

1 to zfs_vdev_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Maximum initializing I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Sch eduler

zfs_vdev_initializing_min_active

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Range:

1 to zfs_vdev_initializing_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Minimum initializing I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Sch eduler

zfs_vdev_max_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1000

Range:

sum of each queue’s min_active to UINT32_MAX

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

The maximum number of I/O operations active to each device. Ideally, this will be at least the sum of each queue’s max_active. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

Notes: The maximum number of I/Os active to each device. Ideally, zfs_vdev_max_active >= the sum of each queue’s max_active. Once queued to the device, the ZFS I/O scheduler is no longer able to prioritize I/O operations. The underlying device drivers have their own scheduler and queue depth limits. Values larger than the device’s maximum queue depth can have the affect of increased latency as the I/Os are queued in the intervening device driver layers.

zfs_vdev_max_auto_ashift

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.1 - master

14

v2.0

ASHIFT_MAX (16)

Change:

Dynamic

Tags:

vdev

Maximum ashift used when optimizing for logical → physical sector size on new top-level vdevs. May be increased up to ASHIFT_MAX (16), but this may negatively impact pool space efficiency.

zfs_vdev_max_ms_shift

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

34

Change:

Dynamic

Tags:

vdev

Default upper limit for metaslab size.

zfs_vdev_min_auto_ashift

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

ASHIFT_MIN

Change:

Dynamic

Tags:

vdev

Minimum ashift used when creating new top-level vdevs.

zfs_vdev_min_ms_count

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

16

Range:

16 to zfs_vdev_ms_count_limit

Change:

Dynamic

Tags:

metaslab, vdev

Minimum number of metaslabs to create in a top-level vdev.

zfs_vdev_mirror_non_rotating_inc

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

mirror, SSD, vdev, vdev_mirror

A number by which the balancing algorithm increments the load calculation for the purpose of selecting the least busy mirror member on non-rotational vdevs when I/O operations do not immediately follow one another.

Notes: The mirror read algorithm uses current load and an incremental weighting value to determine the vdev to service a read operation. Lower values determine the preferred vdev. The weighting value is zfs_vdev_mirror_rotating_inc for rotating media and zfs_vdev_mirror_non_rotating_inc for nonrotating media. Verify the rotational setting described by a block device in sysfs by observing /sys/block/DISK_NAME/queue/rotational

zfs_vdev_mirror_non_rotating_seek_inc

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

mirror, SSD, vdev, vdev_mirror

A number by which the balancing algorithm increments the load calculation for the purpose of selecting the least busy mirror member when an I/O operation lacks locality as defined by the zfs_vdev_mirror_rotating_seek_offset. Operations within this that are not immediately following the previous operation are incremented by half.

Notes: For nonrotating media in a mirror, a seek penalty is applied as sequential I/O’s can be aggregated into fewer operations, avoiding unnecessary per-command overhead, often boosting performance. Verify the rotational setting described by a block device in SysFS by observing /sys/block/DISK_NAME/queue/rotational

zfs_vdev_mirror_rotating_inc

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

HDD, mirror, vdev, vdev_mirror

A number by which the balancing algorithm increments the load calculation for the purpose of selecting the least busy mirror member when an I/O operation immediately follows its predecessor on rotational vdevs for the purpose of making decisions based on load.

When to change: Increasing for mirrors with both rotating and nonrotating media more strongly favors the nonrotating media

Notes: The mirror read algorithm uses current load and an incremental weighting value to determine the vdev to service a read operation. Lower values determine the preferred vdev. The weighting value is zfs_vdev_mirror_rotating_inc for rotating media and zfs_vdev_mirror_non_rotating_inc for nonrotating media. Verify the rotational setting described by a block device in sysfs by observing /sys/block/DISK_NAME/queue/rotational

zfs_vdev_mirror_rotating_seek_inc

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

5

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

HDD, mirror, vdev, vdev_mirror

A number by which the balancing algorithm increments the load calculation for the purpose of selecting the least busy mirror member when an I/O operation lacks locality as defined by zfs_vdev_mirror_rotating_seek_offset. Operations within this that are not immediately following the previous operation are incremented by half.

Notes: For rotating media in a mirror, if the next I/O offset is within zfs_vdev_mirror_rotating_seek_offset then the weighting factor is incremented by (zfs_vdev_mirror_rotating_seek_inc / 2). Otherwise the weighting factor is increased by zfs_vdev_mirror_rotating_seek_inc. This algorithm prefers rotating media with lower seek distance. Verify the rotational setting described by a block device in sysfs by observing /sys/block/DISK_NAME/queue/rotational

zfs_vdev_mirror_rotating_seek_offset

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1048576

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

HDD, mirror, vdev, vdev_mirror

The maximum distance for the last queued I/O operation in which the balancing algorithm considers an operation to have locality. See ZFS I/O SCHEDULER.

Notes: For rotating media in a mirror, if the next I/O offset is within zfs_vdev_mirror_rotating_seek_offset then the weighting factor is incremented by (zfs_vdev_mirror_rotating_seek_inc/ 2). Otherwise the weighting factor is increased by zfs_vdev_mirror_rotating_seek_inc. This algorithm prefers rotating media with lower seek distance. Verify the rotational setting described by a block device in sysfs by observing /sys/block/DISK_NAME/queue/rotational

zfs_vdev_mirror_switch_us

Versions:

v0.6

Platforms:

Linux, FreeBSD

Type:

int

Default:

10,000

Change:

Dynamic

Tags:

vdev, vdev_mirror

Switch mirrors every N usecs

Default value: 10,000.

zfs_vdev_ms_count_limit

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

131072

Range:

zfs_vdev_min_ms_count to 131,072

Change:

Dynamic

Tags:

metaslab, vdev

Practical upper limit of total metaslabs per top-level vdev.

zfs_vdev_nia_credit

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

5

Range:

1 to UINT_MAX

Change:

Dynamic

Tags:

resilver, scrub, vdev, vdev_queue

Some HDDs tend to prioritize sequential I/O so strongly, that concurrent random I/O latency reaches several seconds. On some HDDs this happens even if sequential I/O operations are submitted one at a time, and so setting zfs_*_max_active= 1 does not help. To prevent non-interactive I/O, like scrub, from monopolizing the device, no more than zfs_vdev_nia_credit operations can be sent while there are outstanding incomplete interactive operations. This enforced wait ensures the HDD services the interactive I/O within a reasonable amount of time. See ZFS I/O SCHEDULER.

When to change: See ZIO SCHEDULER

Notes: Some HDDs tend to prioritize sequential I/O so strongly, that concurrent random I/O latency reaches several seconds. On some HDDs this happens even if sequential I/O operations are submitted one at a time, and so setting zfs_*_max_active= 1 does not help. To prevent non-interactive I/O, like scrub, from monopolizing the device, no more than zfs_vdev_nia_credit operations can be sent while there are outstanding incomplete interactive operations. This enforced wait ensures the HDD services the interactive I/O within a reasonable amount of time. See ZIO SCHEDULER.

zfs_vdev_nia_delay

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

5

Range:

1 to UINT_MAX

Change:

Dynamic

Tags:

delay, resilver, scrub, vdev, vdev_queue

For non-interactive I/O (scrub, resilver, removal, initialize and rebuild), the number of concurrently-active I/O operations is limited to zfs_*_min_active, unless the vdev is “idle”. When there are no interactive I/O operations active (synchronous or otherwise), and zfs_vdev_nia_delay operations have completed since the last interactive operation, then the vdev is considered to be “idle”, and the number of concurrently-active non-interactive operations is increased to zfs_*_max_active. See ZFS I/O SCHEDULER.

When to change: See ZIO SCHEDULER

Notes: For non-interactive I/O (scrub, resilver, removal, initialize and rebuild), the number of concurrently-active I/O operations is limited to zfs_*_min_active, unless the vdev is “idle”. When there are no interactive I/O operations active (synchronous or otherwise), and zfs_vdev_nia_delay operations have completed since the last interactive operation, then the vdev is considered to be “idle”, and the number of concurrently-active non-interactive operations is increased to zfs_*_max_active. See ZIO SCHEDULER.

zfs_vdev_open_timeout_ms

Versions:

v2.1 - master

Platforms:

Linux

Type:

uint

Default:

1000

Change:

Dynamic

Tags:

vdev, vdev_disk

Timeout value to wait before determining a device is missing during import. This is helpful for transient missing paths due to links being briefly removed and recreated in response to udev events.

zfs_vdev_queue_depth_pct

Versions:

v0.7 - v2.3

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1000

Range:

1 to UINT32_MAX

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Maximum number of queued allocations per top-level vdev expressed as a percentage of zfs_vdev_async_write_max_active, which allows the system to detect devices that are more capable of handling allocations and to allocate more blocks to those devices. This allows for dynamic allocation distribution when devices are imbalanced, as fuller devices will tend to be slower than empty devices.

Also see zio_dva_throttle_enabled.

When to change: See ZFS I/O Scheduler

Notes: Maximum number of queued allocations per top-level vdev expressed as a percentage of zfs_vdev_async_write_max_active. This allows the system to detect devices that are more capable of handling allocations and to allocate more blocks to those devices. It also allows for dynamic allocation distribution when devices are imbalanced as fuller devices will tend to be slower than empty devices. Once the queue depth reaches (zfs_vdev_queue_depth_pct * zfs_vdev_async_write_max_active / 100) then allocator will stop allocating blocks on that top-level device and switch to the next. See also zio_dva_throttle_enabled

zfs_vdev_raidz_impl

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

string

Default:

fastest

Range:

fastest

fastest implementation selected by microbenchmark

original

original raidz implementation

scalar

scalar raidz implementation

sse2

uses the SSE2 instruction set, 64-bit x86

ssse3

uses the SSSE3 instruction set, 64-bit x86

avx2

uses the AVX2 instruction set, 64-bit x86

avx512f

uses the AVX512F instruction set, 64-bit x86

avx512bw

uses the AVX512F and AVX512BW instruction sets, 64-bit x86

aarch64_neon

uses NEON, aarch64/64 bit ARMv8

aarch64_neonx2

uses NEON with more unrolling, aarch64/64 bit ARMv8

Change:

Dynamic

Tags:

CPU, raidz, vdev

Select the raidz parity implementation to use.

Variants that don’t depend on CPU-specific features may be selected on module load, as they are supported on all systems. The remaining options may only be set after the module is loaded, as they are available only if the implementations are compiled in and supported on the running system.

Once the module is loaded, /sys/module/zfs/parameters/zfs_vdev_raidz_impl will show the available options, with the currently selected one enclosed in square brackets.

fastest

selected by built-in benchmark

original

original implementation

scalar

scalar implementation

sse2

SSE2 instruction set

64-bit x86

ssse3

SSSE3 instruction set

64-bit x86

avx2

AVX2 instruction set

64-bit x86

avx512f

AVX512F instruction set

64-bit x86

avx512bw

AVX512F & AVX512BW instruction sets

64-bit x86

aarch64_neon

NEON

Aarch64/64-bit ARMv8

aarch64_neonx2

NEON with more unrolling

Aarch64/64-bit ARMv8

powerpc_altivec

Altivec

PowerPC

When to change: testing raidz algorithms

Notes: zfs_vdev_raidz_impl overrides the raidz parity algorithm. By default, the algorithm is selected at zfs module load time by the results of a microbenchmark of algorithms based on the current hardware. Once the module is loaded, the content of /sys/module/zfs/parameters/zfs_vdev_raidz_impl shows available options with the currently selected enclosed in []. Details of the results of the microbenchmark are observable in the /proc/spl/kstat/zfs/vdev_raidz_bench file.

zfs_vdev_read_gap_limit

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

32768

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Aggregate read I/O operations if the on-disk gap between them is within this threshold.

Notes: To reduce IOPs, small, adjacent I/Os are aggregated (coalesced) into into a large I/O. For reads, aggregations occur across small adjacency gaps where the gap is less than zfs_vdev_read_gap_limit

zfs_vdev_rebuild_max_active

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

3

Change:

Dynamic

Tags:

vdev, vdev_queue

Maximum sequential resilver I/O operations active to each device. See ZFS I/O SCHEDULER.

zfs_vdev_rebuild_min_active

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Change:

Dynamic

Tags:

vdev, vdev_queue

Minimum sequential resilver I/O operations active to each device. See ZFS I/O SCHEDULER.

zfs_vdev_removal_max_active

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

2

Range:

1 to zfs_vdev_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, vdev_removal, ZIO_scheduler

Maximum removal I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

zfs_vdev_removal_min_active

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Range:

1 to zfs_vdev_removal_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, vdev_removal, ZIO_scheduler

Minimum removal I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

zfs_vdev_scheduler

Versions:

v0.6 - v2.3

Platforms:

Linux

Type:

charp

Range:

expected: noop, cfq, bfq, and deadline

Change:

Dynamic

Tags:

vdev, vdev_disk, ZIO_scheduler

DEPRECATED. Prints warning to kernel log for compatibility.

When to change: since ZFS has its own I/O scheduler, using a simple scheduler can result in more consistent performance

Notes: Prior to version 0.8.3, when the pool is imported, for whole disk vdevs, the block device I/O scheduler is set to zfs_vdev_scheduler. The most common schedulers are: noop, cfq, bfq, and deadline. In some cases, the scheduler is not changeable using this method. Known schedulers that cannot be changed are: scsi_mq and none. In these cases, the scheduler is unchanged and an error message can be reported to logs. The parameter was disabled in v0.8.3 but left in place to avoid breaking loading of the zfs module if the parameter is specified in modprobe configuration on existing installations. It is recommended that users leave the default scheduler “unless you’re encountering a specific problem, or have clearly measured a performance improvement for your workload,” and if so, to change it via the /sys/block/<device>/queue/scheduler interface and/or udev rule.

zfs_vdev_scrub_max_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

2

Range:

1 to zfs_vdev_max_active

Change:

Dynamic

Tags:

resilver, scrub, vdev, vdev_queue, ZIO_scheduler

Maximum scrub I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

Notes: zfs_vdev_scrub_max_active sets the maximum scrub or scan read I/Os active to each device.

zfs_vdev_scrub_min_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Range:

1 to zfs_vdev_scrub_max_active

Change:

Dynamic

Tags:

resilver, scrub, vdev, vdev_queue, ZIO_scheduler

Minimum scrub I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

Notes: zfs_vdev_scrub_min_active sets the minimum scrub or scan read I/Os active to each device.

zfs_vdev_standard_sm_blksz

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Default:

131072

Change:

Dynamic

Tags:

vdev

vdev-wide space maps that have lots of entries written to them at the end of each transaction can benefit from a higher I/O bandwidth (e.g. vdev_obsolete_sm), thus we default their block size to 128K.

zfs_vdev_sync_read_max_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Range:

1 to zfs_vdev_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Maximum synchronous read I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

zfs_vdev_sync_read_min_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Range:

1 to zfs_vdev_sync_read_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Minimum synchronous read I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

zfs_vdev_sync_write_max_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Range:

1 to zfs_vdev_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Maximum synchronous write I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

zfs_vdev_sync_write_min_active

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

10

Range:

1 to zfs_vdev_sync_write_max_active

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Minimum synchronous write I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

zfs_vdev_trim_max_active

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

2

Range:

1 to zfs_vdev_max_active

Change:

Dynamic

Tags:

trim, vdev, vdev_queue, ZIO_scheduler

Maximum trim/discard I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

Notes: zfs_vdev_trim_max_active sets the maximum trim I/Os active to each device.

zfs_vdev_trim_min_active

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Range:

1 to zfs_vdev_trim_max_active

Change:

Dynamic

Tags:

trim, vdev, vdev_queue, ZIO_scheduler

Minimum trim/discard I/O operations active to each device. See ZFS I/O SCHEDULER.

When to change: See ZFS I/O Scheduler

Notes: zfs_vdev_trim_min_active sets the minimum trim I/Os active to each device.

zfs_vdev_write_gap_limit

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

4096

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

vdev, vdev_queue, ZIO_scheduler

Aggregate write I/O operations if the on-disk gap between them is within this threshold.

Notes: To reduce IOPs, small, adjacent I/Os are aggregated (coalesced) into into a large I/O. For writes, aggregations occur across small adjacency gaps where the gap is less than zfs_vdev_write_gap_limit

zfs_vnops_read_chunk_size

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.3 - master

33554432

v2.1 - v2.2

1048576

Change:

Dynamic

Tags:

vnops

Bytes to read per chunk.

zfs_wrlog_data_max

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

u64

Change:

Dynamic

Tags:

DSL, dsl_pool

The upper limit of write-transaction ZIL log data size in bytes. Write operations are throttled when approaching the limit until log data is cleared out after transaction group sync. Because of some overhead, it should be set at least 2 times the size of zfs_dirty_data_max to prevent harming normal write throughput. It also should be smaller than the size of the slog device if slog is present.

Defaults to zfs_dirty_data_max*2

zfs_xattr_compat

Versions:

v2.2 - master

Platforms:

Linux

Type:

int

Range:

0 | 1

Change:

Dynamic

Tags:

xattr

Control the naming scheme used when setting new xattrs in the user namespace. If 0 (the default on Linux), user namespace xattr names are prefixed with the namespace, to be backwards compatible with previous versions of ZFS on Linux. If 1 (the default on FreeBSD), user namespace xattr names are not prefixed, to be backwards compatible with previous versions of ZFS on illumos and FreeBSD.

Either naming scheme can be read on this and future versions of ZFS, regardless of this tunable, but legacy ZFS on illumos or FreeBSD are unable to read user namespace xattrs written in the Linux format, and legacy versions of ZFS on Linux are unable to read user namespace xattrs written in the legacy ZFS format.

An existing xattr with the alternate naming scheme is removed when overwriting the xattr so as to not accumulate duplicates.

zfs_zevent_cols

Versions:

v0.6 - v2.0

Platforms:

Linux, FreeBSD

Type:

int

Default:

80

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

debug, fm, zed

When zevents are logged to the console use this as the word wrap width.

Default value: 80.

When to change: if 80 columns isn’t enough

Notes: zfs_zevent_cols is a soft wrap limit in columns (characters) for ZFS events logged to the console.

zfs_zevent_console

Versions:

v0.6 - v2.0

Platforms:

Linux, FreeBSD

Type:

int

Range:

0

do not log to console

1

log to console

Change:

Dynamic

Tags:

debug, fm, zed

Log events to the console

Use 1 for yes and 0 for no (default).

When to change: to log ZFS events to the console

Notes: If zfs_zevent_console is true (1), then ZFS events are logged to the console. More logging and log filtering capabilities are provided by zed

zfs_zevent_len_max

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.0 - master

512

v0.6 - v0.8

0

Range:

0 to INT_MAX

Change:

Dynamic

Tags:

debug, fm, zed

Max event queue length. Events in the queue can be viewed with zpool-events(8).

When to change: increase to see more ZFS events

Notes: zfs_zevent_len_max is the maximum ZFS event queue length. A value of 0 results in a calculated value (16 * number of CPUs) with a minimum of 64. Events in the queue can be viewed with the zpool events command.

zfs_zevent_retain_expire_secs

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

900

Change:

Dynamic

Tags:

fm, zed

Lifespan for a recent ereport that was retained for duplicate checking.

zfs_zevent_retain_max

Versions:

v2.0 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

2000

Change:

Dynamic

Tags:

fm, zed

Maximum recent zevent records to retain for duplicate checking. Setting this to 0 disables duplicate detection.

zfs_zil_clean_taskq_maxalloc

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1048576

Range:

zfs_zil_clean_taskq_minalloc to INT_MAX

Change:

Dynamic

Tags:

DSL, dsl_pool, taskq, ZIL

The maximum number of taskq entries that are allowed to be cached. When this limit is exceeded transaction records (itxs) will be cleaned synchronously.

When to change: If more dp_zil_clean_taskq entries are needed to prevent the itxs from being synchronously cleaned

Notes: During a SPA sync, intent log transaction groups (itxg) are cleaned. The cleaning work is dispatched to the DSL pool ZIL clean taskq (dp_zil_clean_taskq). zfs_zil_clean_taskq_minalloc is the minimum and zfs_zil_clean_taskq_maxalloc is the maximum number of cached taskq entries for dp_zil_clean_taskq. The actual number of taskq entries dynamically varies between these values. When zfs_zil_clean_taskq_maxalloc is exceeded transaction records (itxs) are cleaned synchronously with possible negative impact to the performance of SPA sync. Ideally taskq entries are pre-allocated prior to being needed by zil_clean(), thus avoiding dynamic allocation of new taskq entries.

zfs_zil_clean_taskq_minalloc

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1024

Range:

1 to zfs_zil_clean_taskq_maxalloc

Change:

Dynamic

Tags:

DSL, dsl_pool, taskq, ZIL

The number of taskq entries that are pre-populated when the taskq is first created and are immediately available for use.

Notes: During a SPA sync, intent log transaction groups (itxg) are cleaned. The cleaning work is dispatched to the DSL pool ZIL clean taskq (dp_zil_clean_taskq). zfs_zil_clean_taskq_minalloc is the minimum and zfs_zil_clean_taskq_maxalloc is the maximum number of cached taskq entries for dp_zil_clean_taskq. The actual number of taskq entries dynamically varies between these values. zfs_zil_clean_taskq_minalloc is the minimum number of ZIL transaction records (itxs). Ideally taskq entries are pre-allocated prior to being needed by zil_clean(), thus avoiding dynamic allocation of new taskq entries.

zfs_zil_clean_taskq_nthr_pct

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

100

Range:

1 to 100

Change:

Dynamic

Tags:

DSL, dsl_pool, taskq, ZIL

This controls the number of threads used by dp_zil_clean_taskq. The default value of 100% will create a maximum of one thread per CPU.

When to change: Testing ZIL clean and SPA sync performance

Notes: zfs_zil_clean_taskq_nthr_pct controls the number of threads used by the DSL pool ZIL clean taskq (dp_zil_clean_taskq). The default value of 100% will create a maximum of one thread per cpu.

zfs_zil_saxattr

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

ZIL

Setting this tunable to zero disables ZIL logging of new xattr=sa records if the org.openzfs:zilsaxattr feature is enabled on the pool. This would only be necessary to work around bugs in the ZIL logging or replay code for this record type. The tunable has no effect if the feature is disabled.

zil_maxblocksize

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

131072

Change:

Dynamic

Tags:

ZIL

This sets the maximum block size used by the ZIL. On very fragmented pools, lowering this (typically to 36 KiB) can improve performance.

zil_maxcopied

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

7680

Change:

Dynamic

Tags:

ZIL

This sets the maximum number of write bytes logged via WR_COPIED. It tunes a tradeoff between additional memory copy and possibly worse log space efficiency vs additional range lock/unlock.

zil_min_commit_timeout

Versions:

v2.1

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

5000

Change:

Dynamic

Tags:

ZIL

This sets the minimum delay in nanoseconds ZIL care to delay block commit, waiting for more records. If ZIL writes are too fast, kernel may not be able sleep for so short interval, increasing log latency above allowed by zfs_commit_timeout_pct.

zil_nocacheflush

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

send cache flush commands

1

do not send cache flush commands

Change:

Dynamic

Tags:

disks, ZIL

Disable the cache flush commands that are normally sent to disk by the ZIL after an LWB write has completed. Setting this will cause ZIL corruption on power loss if a volatile out-of-order write cache is enabled.

When to change: If the storage device has nonvolatile cache, then disabling cache flush can save the cost of occasional cache flush commands

Notes: ZFS uses barriers (volatile cache flush commands) to ensure data is committed to permanent media by devices. This ensures consistent on-media state for devices where caches are volatile (eg HDDs). zil_nocacheflush disables the cache flush commands that are normally sent to devices by the ZIL after a log write has completed. The difference between zil_nocacheflush and zfs_nocacheflush is zil_nocacheflush applies to ZIL writes while zfs_nocacheflush disables barrier writes to the pool devices at the end of transaction group syncs. WARNING: setting this can cause ZIL corruption on power loss if the device has a volatile write cache.

zil_replay_disable

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

replay ZIL

1

destroy ZIL

Change:

Dynamic

Tags:

debug, ZIL

Disable intent logging replay. Can be disabled for recovery from corrupted ZIL.

When to change: Do not change

Notes: If zil_replay_disable = 1, then when a volume or filesystem is brought online, no attempt to replay the ZIL is made and any existing ZIL is destroyed. This can result in loss of data without notice.

zil_slog_bulk

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

u64

Default:

v2.2 - master

67108864

v0.7 - v2.1

786,432

Range:

0 to ULONG_MAX

Change:

Dynamic

Tags:

ZIL

Limit SLOG write size per commit executed with synchronous priority. Any writes above that will be executed with lower (asynchronous) priority to limit potential SLOG device abuse by single active ZIL writer.

When to change: See ZFS I/O Scheduler

Notes: zil_slog_bulk is the log device write size limit per commit executed with synchronous priority. Writes below zil_slog_bulk are executed with synchronous priority. Writes above zil_slog_bulk are executed with lower (asynchronous) priority to reduct potential log device abuse by a single active ZIL writer.

zil_slog_limit

Versions:

v0.6

Platforms:

Linux, FreeBSD

Type:

ulong

Default:

1,048,576

Change:

Dynamic

Tags:

ZIL

Max commit bytes to separate log device

Default value: 1,048,576.

zil_special_is_slog

Versions:

v2.4 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

ZIL

When enabled, and written blocks go to normal vdevs, treat present special vdevs as SLOGs. Blocks that go to the special vdevs are still written indirectly, as with logbias=throughput. This parameter is ignored if an SLOG is present.

zio_batch_enabled

Versions:

master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

1 | 0

Change:

Dynamic

Tags:

ZIO

Collect the vdev children of a RAIDZ, dRAID or mirror I/O operation as they complete, and process all of their completions on a single thread once the last of them is in, rather than one thread each. Applies only to children that bypass the vdev I/O queue.

zio_deadman_log_all

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

do not log all deadman events

1

log all deadman events

Change:

Dynamic

Tags:

debug, ZIO

If non-zero, the zio deadman will produce debugging messages (see zfs_dbgmsg_enable) for all zios, rather than only for leaf zios possessing a vdev. This is meant to be used by developers to gain diagnostic information for hang conditions which don’t involve a mutex or other locking primitive: typically conditions in which a thread in the zio pipeline is looping indefinitely.

When to change: when debugging ZFS I/O pipeline

Notes: zio_deadman_log_all enables debugging messages for all ZFS I/Os, rather than only for leaf ZFS I/Os for a vdev. This is meant to be used by developers to gain diagnostic information for hang conditions which don’t involve a mutex or other locking primitive. Typically these are conditions where a thread in the zio pipeline is looping indefinitely. See also zfs_dbgmsg_enable

zio_delay_max

Versions:

v0.6 - v0.7

Platforms:

Linux, FreeBSD

Type:

int

Default:

30,000

Range:

1 to INT_MAX

Change:

Dynamic

Tags:

debug, delay, ZIO

A zevent will be logged if a ZIO operation takes more than N milliseconds to complete. Note that this is only a logging facility, not a timeout on operations.

Default value: 30,000.

When to change: when debugging slow I/O

Notes: If a ZFS I/O operation takes more than zio_delay_max milliseconds to complete, then an event is logged. Note that this is only a logging facility, not a timeout on operations. See also zpool events

zio_dva_throttle_enabled

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

1

Range:

0

do not throttle ZIO block allocations

1

throttle ZIO block allocations

Change:

Dynamic

Tags:

vdev, ZIO, ZIO_scheduler

Throttle block allocations in the I/O pipeline. This allows for dynamic allocation distribution based on device performance.

When to change: Testing ZIO block allocation algorithms

Notes: zio_dva_throttle_enabled controls throttling of block allocations in the ZFS I/O (ZIO) pipeline. When enabled, the maximum number of pending allocations per top-level vdev is limited by zfs_vdev_queue_depth_pct (removed after v2.3)

zio_requeue_io_start_cut_in_line

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0

don’t prioritize re-queued I/Os

1

prioritize re-queued I/Os

Change:

Dynamic

Tags:

ZIO, ZIO_scheduler

Prioritize requeued I/O.

When to change: Do not change

Notes: zio_requeue_io_start_cut_in_line controls prioritization of a re-queued ZFS I/O (ZIO) in the ZIO pipeline by the ZIO taskq.

zio_slow_io_ms

Versions:

v0.8 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

30000

Range:

0 to MAX_INT

Change:

Dynamic

Tags:

vdev, zed, ZIO

When an I/O operation takes more than this much time to complete, it’s marked as slow. Each slow operation causes a delay zevent. Slow I/O counters can be seen with zpool status -s.

When to change: when debugging slow devices and the default value is inappropriate

Notes: An I/O operation taking more than zio_slow_io_ms milliseconds to complete is marked as a slow I/O. Slow I/O counters can be observed with zpool status -s. Each slow I/O causes a delay zevent, observable using zpool events. See also zfs-events(5).

zio_taskq_batch_pct

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.1 - master

80

v0.7 - v2.0

75

Range:

1 to 100, fractional number of CPUs are rounded down

Change:

Dynamic

Tags:

SPA, taskq, ZIO, ZIO_scheduler

Percentage of online CPUs which will run a worker thread for I/O. These workers are responsible for I/O work such as compression, encryption, checksum and parity calculations. Fractional number of CPUs will be rounded down.

The default value of 80% was chosen to avoid using all CPUs which can result in latency issues and inconsistent application performance, especially when slower compression and/or checksumming is enabled. Set value only applies to pools imported/created after that.

When to change: To tune parallelism in multiprocessor systems

Verification: The number of taskqs for each batch group can be observed using ps and counting the threads

Notes: zio_taskq_batch_pct sets the number of I/O worker threads as a percentage of online CPUs. These workers threads are responsible for IO work such as compression and checksum calculations. Each block is handled by one worker thread, so maximum overall worker thread throughput is function of the number of concurrent blocks being processed, the number of worker threads, and the algorithms used. The default is chosen to avoid using all CPUs which can result in latency issues and inconsistent application performance, especially when high compression is enabled. The taskq batch processes are:

zio_taskq_batch_tpq

Versions:

v2.1 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Change:

Dynamic

Tags:

SPA, taskq, ZIO

Number of worker threads per taskq. Higher values improve I/O ordering and CPU utilization, while lower reduce lock contention. Set value only applies to pools imported/created after that.

If 0, generate a system-dependent value close to 6 threads per taskq. Set value only applies to pools imported/created after that.

zio_taskq_write_tpq

Versions:

v2.3 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

16

Change:

Dynamic

Tags:

SPA, taskq, ZIO

Determines the minimum number of threads per write issue taskq. Higher values improve CPU utilization on high throughput, while lower reduce taskq locks contention on high IOPS. Set value only applies to pools imported/created after that.

zstd_abort_size

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

131072

Change:

Dynamic

Tags:

zstd

Minimal uncompressed size (inclusive) of a record before the early abort heuristic will be attempted.

zstd_earlyabort_pass

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Change:

Dynamic

Tags:

zstd

Whether heuristic for detection of incompressible data with zstd levels >= 3 using LZ4 and zstd-1 passes is enabled.

zvol_blk_mq_blocks_per_thread

Versions:

v2.2 - master

Platforms:

Linux

Type:

uint

Default:

8

Change:

Dynamic

Tags:

ZVOL

If zvol_use_blk_mq is enabled, then process this number of volblocksize-sized blocks per zvol thread. This tunable can be use to favor better performance for zvol reads (lower values) or writes (higher values). If set to 0, then the zvol layer will process the maximum number of blocks per thread that it can. This parameter will only appear if your kernel supports blk-mq and is only applied at each zvol’s load time.

zvol_blk_mq_queue_depth

Versions:

v2.2 - master

Platforms:

Linux

Type:

uint

Default:

0

Change:

Dynamic

Tags:

ZVOL

The queue_depth value for the zvol blk-mq interface. This parameter will only appear if your kernel supports blk-mq and is only applied at each zvol’s load time. If 0 (the default) then use the kernel’s default queue depth. Values are clamped to the kernel’s BLKDEV_MIN_RQ and BLKDEV_MAX_RQ/BLKDEV_DEFAULT_RQ limits.

zvol_enforce_quotas

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

int

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

DSL

Enable strict ZVOL quota enforcement. The strict quota enforcement may have a performance impact.

zvol_inhibit_dev

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0

create volume device nodes

1

do not create volume device nodes

Change:

Dynamic

Tags:

import, volume, ZVOL

Do not create zvol device nodes. This may slightly improve startup time on systems with a very large number of zvols.

When to change: Inhibiting can slightly improve startup time on systems with a very large number of volumes

Notes: zvol_inhibit_dev controls the creation of volume device nodes upon pool import.

zvol_major

Versions:

v0.6 - master

Platforms:

Linux

Type:

uint

Default:

230

Change:

Prior to module load

Tags:

volume, ZVOL

Major number for zvol block devices.

When to change: Do not change

zvol_max_discard_blocks

Versions:

v0.6 - master

Platforms:

Linux

Type:

ulong

Default:

16384

Change:

Prior to module load

Tags:

discard, volume, ZVOL

Discard (TRIM) operations done on zvols will be done in batches of this many blocks, where block size is determined by the volblocksize property of a zvol.

When to change: if volume discard activity severely impacts other workloads

Verification: Observe value of /sys/block/ VOLUME_INSTANCE/queue/discard_max_bytes

Notes: Discard (aka ATA TRIM or SCSI UNMAP) operations done on volumes are done in batches zvol_max_discard_blocks blocks. The block size is determined by the volblocksize property of a volume. Some applications, such as mkfs, discard the whole volume at once using the maximum possible discard size. As a result, many gigabytes of discard requests are not uncommon. Unfortunately, if a large amount of data is already allocated in the volume, ZFS can be quite slow to process discard requests. This is especially true if the volblocksize is small (eg default=8KB). As a result, very large discard requests can take a very long time (perhaps minutes under heavy load) to complete. This can cause a number of problems, most notably if the volume is accessed remotely (eg via iSCSI), in which case the client has a high probability of timing out on the request. Limiting the zvol_max_discard_blocks can decrease the amount of discard workload request by setting the discard_max_bytes and discard_max_hw_bytes for the volume’s block device in SysFS. This value is readable by volume device consumers.

zvol_num_taskqs

Versions:

v2.2 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Change:

Dynamic

Tags:

ZVOL

Number of zvol taskqs. If 0 (the default) then scaling is done internally to prefer 6 threads per taskq. This only applies on Linux.

zvol_open_timeout_ms

Versions:

v2.2 - master

Platforms:

Linux

Type:

uint

Change:

Dynamic

Tags:

ZVOL

Timeout for ZVOL open retries

zvol_prefetch_bytes

Versions:

v0.6 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

131072

Change:

Dynamic

Tags:

prefetch, volume, ZVOL

When adding a zvol to the system, prefetch this many bytes from the start and end of the volume. Prefetching these regions of the volume is desirable, because they are likely to be accessed immediately by blkid(8) or the kernel partitioner.

Notes: When importing a pool with volumes or adding a volume to a pool, zvol_prefetch_bytes are prefetch from the start and end of the volume. Prefetching these regions of the volume is desirable because they are likely to be accessed immediately by blkid(8) or by the kernel scanning for a partition table.

zvol_request_sync

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

0

Range:

0

do concurrent (async) volume requests

1

do sync volume requests

Change:

Dynamic

Tags:

volume, ZVOL

When processing I/O requests for a zvol, submit them synchronously. This effectively limits the queue depth to 1 for each I/O submitter. When unset, requests are handled asynchronously by a thread pool. The number of requests which can be handled concurrently is controlled by zvol_threads. This parameter applies regardless of the zvol_use_blk_mq setting.

When to change: Testing concurrent volume requests

Notes: When processing I/O requests for a volume submit them synchronously. This effectively limits the queue depth to 1 for each I/O submitter. When set to 0 requests are handled asynchronously by the “zvol” thread pool. See also zvol_threads

zvol_threads

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

v2.2 - master

0

v0.7 - v2.1

32

Range:

0=one thread per active CPU, otherwise 1 to UINT_MAX

Change:

Dynamic

Tags:

volume, ZVOL

The number of system wide threads to use for processing zvol block IOs. If 0 (the default) then internally set zvol_threads to the number of CPUs present or 32 (whichever is greater).

When to change: Matching the number of concurrent volume requests with workload requirements can improve concurrency

Verification: iostat using avgqu-sz or aqu-sz results

Notes: zvol_threads controls the maximum number of threads handling concurrent volume I/O requests. Since v2.2.0 the default of 0 means one thread per active CPU; before that it was 32, which behaves similarly to a disk with a 32-entry command queue. The actual number of threads required can vary widely by workload and available CPUs. If lock analysis shows high contention in the zvol taskq threads, then reducing the number of zvol_threads or workload queue depth can improve overall throughput. See also zvol_request_sync.

zvol_use_blk_mq

Versions:

v2.2 - master

Platforms:

Linux

Type:

uint

Default:

0

Range:

0 | 1

Change:

Dynamic

Tags:

ZVOL

Set to 1 to use the blk-mq API for zvols. Set to 0 (the default) to use the legacy zvol APIs. This setting can give better or worse zvol performance depending on the workload. This parameter will only appear if your kernel supports blk-mq and is only read and assigned to a zvol at zvol load time.

zvol_volmode

Versions:

v0.7 - master

Platforms:

Linux, FreeBSD

Type:

uint

Default:

1

Range:

1

full, legacy fully functional behaviour (default)

2

dev, hide partitions on volume block devices

3

none, not exposing volumes outside ZFS

Change:

Dynamic

Tags:

volume, ZVOL

Defines zvol block devices behavior when volmode=default:

1

equivalent to full

2

equivalent to dev

3

equivalent to none

Notes: zvol_volmode defines volume block devices behaviour when the volmode property is set to default. To maintain compatibility with ZFS on BSD, “geom” is synonymous with “full”.