Perception and Hearing
Standards: ISO 226ECMA-418ISO 389ISO 7029ISO 1999Key references: Fletcher & Munson 1933Fastl & Zwicker 2007Houtgast & Steeneken 1985French & Steinberg 1947+1 more
This page collects the theory behind hearing and psychoacoustics: the equal-loudness contours, the Zwicker, Moore-Glasberg and Sottek loudness models, the sound-quality metrics tonality, roughness and sharpness, tone prominence, the speech metrics STI and SII, and the statistics of hearing thresholds and hearing loss. It is part of the theory reference.
Equal-loudness contours (ISO 226:2023)
Section titled “Equal-loudness contours (ISO 226:2023)”A tone has a loudness level of phon when it is judged equally loud as a 1 kHz pure tone at dB SPL. ISO 226:2023 Formula (1) (clause 4.1, p. 2) gives the SPL of a pure tone at frequency that reaches loudness level :
Formula (2) (clause 4.2) inverts it, returning the loudness level of a tone at SPL :
The three parameters come from Table 1 (p. 4), tabulated at the 29 preferred third-octave frequencies of ISO 266 from 20 Hz to 12.5 kHz:
- : exponent for loudness perception at frequency ,
- : magnitude of the linear transfer function, normalized at 1 kHz ( at 1 kHz),
- : threshold of hearing at , in dB.
The standard specifies no interpolation between the tabulated frequencies. Formula (1) is specified for 20 phon to 90 phon between 20 Hz and 4 kHz, and only up to 80 phon between 5 kHz and 12.5 kHz; above 80 phon the contour therefore stops at 4 kHz. Values outside these limits from Formula (2) are extrapolations the standard labels as informative only.
See the Loudness guide for usage.
The ISO 226:2023 contours from Formula (1), 20 to 90 phon, with the hearing threshold.
Tone prominence: TNR and PR (ECMA-418-1)
Section titled “Tone prominence: TNR and PR (ECMA-418-1)”Both methods operate on a Hann-windowed, RMS-averaged power spectrum (clauses 11.1 / 12.1) and use the clause 10 critical-band model. The critical bandwidth centred on a tone at is (Formula 2):
Band edges are placed arithmetically for Hz (Formulae 4–5): , and geometrically above (Formulae 7–8): , .
TNR (clause 11). The tone band spans the spectral minima on both sides of the peak within 15 % of (clause 11.2). The tone power subtracts the straight line connecting the band-edge bins (Formula 9): over tone-band bins, . The masking-noise power is the remaining critical-band power rescaled to the full critical bandwidth (Formula 10): , and (Formula 11). The prominence criterion (Formulae 12–13) is
PR (clause 12) compares the level of the critical band centred on the tone, , with the mean power of the two contiguous critical bands , (edges from the fitted Formulae 21–22 with Tables 2–3): (Formula 23). For Hz the lower band is truncated at 20 Hz and its power rescaled to a 100 Hz bandwidth (Formula 24). The criterion (Formulae 25–26) is 9.0 dB at kHz, rising as below. Tones are assessed within the 89.1 Hz – 11.2 kHz range of interest (clauses 11.5 / 12.6).
See the Prominent Discrete Tones guide for usage.
Zwicker loudness (ISO 532-1)
Section titled “Zwicker loudness (ISO 532-1)”The ear analyzes sound in critical bands: frequency regions within which energy is summed before loudness is formed. The Bark scale maps frequency to critical-band rate , 0 to 24 Bark, and ISO 532-1:2017 samples the specific loudness at 0.1-Bark steps (240 values). The implementation is a clean-room port of the standard’s normative reference program (Annex A.4) and proceeds in stages:
-
One-third-octave levels: 28 bands, 25 Hz to 12.5 kHz (the Annex A filterbank at 48 kHz, Tables A.1/A.2). For time-varying sounds the squared band outputs are smoothed by three cascaded low-passes with ( capped at 1 kHz) and sampled every 2 ms.
-
Low-frequency grouping: the 11 bands up to 250 Hz receive the equal-loudness corrections of Table A.3 and are summed into the first three critical bands (25–80, 100–160, 200–250 Hz).
-
a0 transmission: the outer/middle-ear transfer correction of Table A.4 (plus the diffuse-field difference of Table A.5 when
field='diffuse') yields the critical-band levels . -
Core loudness: each of the 20 critical bands is transformed with the threshold-in-quiet levels of Table A.6 (after the bandwidth adaptation DCB of Table A.7):
(the reference program’s form of Zwicker’s loudness transformation; bands below threshold contribute zero).
-
Slopes: level-dependent upper masking slopes (steepness per specific-loudness range and critical band, Tables A.8/A.9) attach decaying flanks toward higher ; the total loudness is the area under the pattern:
For time-varying sounds a nonlinear temporal decay (time constants 5/15/75 ms, clause 6.3) and the duration-dependent weighting of the total loudness (3.5 ms and 70 ms low-passes weighted 0.47/0.53, clause 6.4) precede the 500 Hz loudness-vs-time output and the percentile values N5/N10 (clause 6.5).
Sone and phon are tied together by the 1 kHz anchor (1 sone = 40 phon; clause 5.6):
below 1 sone the reference program uses , floored at 3 phon.
See the Loudness guide for usage.
Specific loudness N′(z) over the Bark axis: energy spread over many critical bands sums to more sones than the same band level in a single band.
Advanced loudness models & sound quality
Section titled “Advanced loudness models & sound quality”ISO 532-1 is one of three loudness models; two newer families refine the auditory front-end and add the sound-quality metrics tonality and roughness.
Moore-Glasberg loudness (ISO 532-2:2017, ISO 532-3:2023)
Section titled “Moore-Glasberg loudness (ISO 532-2:2017, ISO 532-3:2023)”Instead of Zwicker’s fixed critical bands, the Moore-Glasberg model forms a continuous excitation pattern on the ERB-number (“Cam”) scale using level-dependent rounded-exponential (roex) auditory filters. As a function of the normalized frequency deviation from a filter centred at , the filter weighting is
where the slope grows with the source level, broadening the lower skirt as level rises (ISO 532-2, Formulae 2–5); this reproduces the upward spread of masking. Passing the stimulus intensity through every filter gives the excitation , and a compressive law maps it to the specific loudness in sone/Cam (Formulae 7–9), of the mid-level form
with the calibration constant sone/Cam (ISO 532-2; in ISO 532-3). The total loudness is the area under the pattern,
and a binaural-inhibition stage (Formulae 10–13) combines the ears so a diotic sound is louder than the same sound at one ear. The 1 kHz / 40 dB SPL anchor gives exactly 1 sone.
ISO 532-3 makes this time-varying. A running spectrum from six parallel Hann-windowed FFTs (segment lengths 2–64 ms, each contributing its own frequency range, updated every ms) drives the same excitation and specific-loudness chain, integrated by two cascaded first-order smoothers with ,
using a fast time constant on the attack and a slower one on the release. This yields the short-term loudness (attack/release near 20–30 ms) and the long-term loudness (near 0.1–0.75 s); the peak long-term loudness predicts the loudness of sounds up to about 5 s.
Sottek Hearing Model (ECMA-418-2:2025)
Section titled “Sottek Hearing Model (ECMA-418-2:2025)”ECMA-418-2 builds all three of its metrics on one auditory front-end (Clause 5): an outer/middle-ear filter, a bank of 53 overlapping gammatone-like band-pass filters spaced on the Bark_HMS scale ( to ), half-wave rectification, and a short-block RMS per band and time block . A compressive nonlinearity (Formula 23) turns the band RMS into the specific basis loudness , whose calibration constant fixes a 1 kHz / 40 dB SPL tone at 1 sone_HMS. The loudness assembles the tonal and noise loudness (below) over bands and time (Formulae 113–117); it grows about per 10 dB, more slowly than Zwicker’s factor of 2, an intrinsic property of the Sottek summation.
Tonality: autocorrelation of the band signal (ECMA-418-2)
Section titled “Tonality: autocorrelation of the band signal (ECMA-418-2)”A tonal component is periodic, so it survives in the autocorrelation function (ACF) of a band’s rectified signal while broadband noise decorrelates. For each band the unbiased ACF of the block is
A windowed spectral estimate of separates a tonal loudness from the noise loudness (Formulae 36–48). The specific tonality is the tonal loudness scaled by a smooth signal-to-noise gate (Formulae 49–51),
and the single value (tu_HMS) is the gated time-average of the per-block maximum over bands (Formulae 61–64). The constant fixes the 1 kHz / 40 dB tone at 1 tu_HMS, and the band of the ACF peak gives the tonal frequency .
Roughness: envelope modulation (ECMA-418-2)
Section titled “Roughness: envelope modulation (ECMA-418-2)”Roughness is the sensation of fast (roughly 20–300 Hz) amplitude modulation, strongest near 70 Hz. From each band’s envelope (Hilbert magnitude), a modulation spectrum is formed and weighted by a modulation-rate function peaking near 70 Hz and by the modulation depth; correlating the modulation across neighbouring bands and applying the specified temporal filtering yields the specific roughness and the time-dependent roughness
(Formulae 65–111). The single value is the 90th percentile of over time (Clause 7.1.10); the constant (Formula 104) calibrates the reference sound (a 1 kHz carrier 100 % amplitude-modulated at 70 Hz at 60 dB SPL) to 1 asper.
Sharpness (DIN 45692)
Section titled “Sharpness (DIN 45692)”Sharpness condenses the high-frequency emphasis of a sound into one number: the -weighted first moment of the ISO 532-1 stationary specific-loudness pattern (DIN 45692:2009, Equation 1):
evaluated on the same 240-bin, 0.1-Bark grid. The constant is not hard-coded but derived from the calibration requirement (clause 6): a critical-band-wide narrowband noise 920–1080 Hz at 60 dB SPL scores exactly 1 acum, and the derived lands inside the normative window (clause 5.2). The informative Annex B weightings are provided under the same 1-acum anchor: von Bismarck (knee at 15 Bark, ) and Aures (loudness-dependent, ). The Table A.2 narrow-band targets are reproduced within the clause 6 tolerance (5 % or 0.05 acum): 0.38 acum at 250 Hz, 1.00 at 1 kHz, 1.78 at 2.5 kHz, 2.82 at 4 kHz.
See the Sound Quality Metrics guide for usage.
Modulation transfer and STI (IEC 60268-16)
Section titled “Modulation transfer and STI (IEC 60268-16)”Speech intelligibility rides on the slow intensity modulations of the speech envelope. The modulation transfer function of a transmission channel is the ratio of received to emitted modulation depth of the octave-band intensity envelope at modulation frequency ; the full STI evaluates it at the 14 one-third-octave modulation frequencies 0.63–12.5 Hz in the seven octave bands 125 Hz – 8 kHz (A.2.2). From a measured impulse response the Schroeder closed form gives it directly (indirect method):
Steady background noise multiplies each band’s by the intensity ratio (the noise term):
and when absolute band levels are known the full correction adds the auditory masking intensity (from the next lower octave band, Table A.2) and the absolute reception threshold (Table A.3). Each corrected maps to an effective SNR, clipped to the ±15 dB range where intelligibility actually varies, then to a transmission index (A.5.4/A.5.5):
The band MTI is the mean TI over the modulation frequencies, and the STI weights the bands with the male factors , of Ed. 5 Table A.1 (A.5.6):
truncated to 1.0. STIPA (Annex B) samples the same physics with just two modulation frequencies per band (Table B.1) on a test signal with source modulation index 0.55; the received depths are measured by sine/cosine correlation of the ~100 Hz low-passed intensity envelopes over an integer number of modulation periods:
See the Speech Transmission Index guide for usage.
Speech Intelligibility Index (ANSI S3.5)
Section titled “Speech Intelligibility Index (ANSI S3.5)”Where the STI characterizes a transmission channel, the SII (ANSI S3.5-1997) predicts intelligibility from what the listener can actually hear: 18 one-third-octave bands 160 Hz – 8 kHz, each contributing its band importance (Table 3, , peaking near 2 kHz). All inputs are equivalent spectrum levels (clauses 3.11/3.55). Speech masks itself upward: each band’s masking spectrum (clause 5.4) accumulates the lower bands along slopes dB, and the disturbance is the larger of masking and hearing floor, (clause 5.6), with the reference internal noise spectrum plus the listener’s hearing-threshold shift (clauses 5.5/5.6). The band audibility clips the speech-to-disturbance margin into (clause 5.8), a level-distortion factor discounts overly loud presentation (clause 5.7), and the index sums (clause 6):
The Table 3 standard speech spectra for the normal, raised, loud and shout vocal efforts are built in (25.01 / 33.86 / 42.16 / 51.31 dB at 1 kHz); in the level-distortion factor is always the normal-effort spectrum. The anchor values: the normal-effort spectrum in quiet with normal hearing scores SII ≈ 0.996, the masking-spectrum reference values are matched to , and the vocal-effort spectra are cross-verified against the Google and CRAN reference implementations.
See the Speech Intelligibility guide for usage.
Hearing thresholds and presbycusis (ISO 389-7, ISO 7029)
Section titled “Hearing thresholds and presbycusis (ISO 389-7, ISO 7029)”ISO 389-7:2005 Table 1 fixes the reference threshold of hearing of otologically normal young adults: the free-field and diffuse-field SPL corresponding to 0 dB HL at the 11 audiometric frequencies 125 Hz – 8 kHz (22.1 dB at 125 Hz for both fields, 2.4/0.8 dB free/diffuse at 1 kHz, diverging at high frequency to 12.6 vs 6.8 dB at 8 kHz). ISO 7029:2017 describes how that threshold shifts statistically with age: the median deviation from age 18 is (clause 4.2, Table 1)
and any fractile follows a two-sided Gaussian model (clause 4.4), , using the upper spread for (worse than median) and the lower spread otherwise, each a degree-5 polynomial in per sex and frequency (clause 4.3, Tables 2–5). At age 18 every deviation is zero by construction. The formulae are established to 80 years at and below 2 kHz and to 70 years above; beyond that the evaluation is an extrapolation. Anchors: at 60 years the medians evaluate to 7.85 dB (male, 1 kHz), 20.21 dB (male, 4 kHz) and 15.32 dB (female, 4 kHz), matching the Table 1 formula to .
See the Hearing Threshold guide for usage.
The ISO 7029 median age shift with its fractile band (left) and the ISO 389-7 free- and diffuse-field reference thresholds (right).
Noise-induced hearing loss (ISO 1999)
Section titled “Noise-induced hearing loss (ISO 1999)”ISO 1999:2013 predicts the permanent threshold shift a noise-exposed population accrues. The median noise-induced shift (NIPTS) for 10–40 years of exposure is (clause 6.3.1, Formula 2, Table 1):
quadratic in the excess over the frequency-dependent onset level (75 dB at 4 kHz, the most sensitive band, up to 93 dB at 500 Hz) and zero below it; under 10 years it scales as (Formula 3). Fractiles add the spread, with (clause 6.3.2, Formulae 4–7, Tables 2/3), clamped at zero. The library’s fractile counts the fraction of the population with the smaller shift, so fractile=0.9 is the most-susceptible decile (reliable range 0.05–0.95); ISO 1999’s own runs the other way, counting the percentage with worse hearing, so the same decile is there and in the Annex D column headings. The hearing threshold level associated with age and noise (HTLAN) combines NIPTS with the ISO 7029 age component at the same fractile through the compressed sum (clause 6.1, Formula 1):
The Annex D worked examples (Tables D.1–D.4; e.g. 100 dB / 40 yr at 3 kHz: 29/38/60 dB at the 0.10/0.50/0.90 fractiles) are reproduced exactly at the standard’s integer rounding, and the Formula 2 hand value at 4 kHz / 20 yr / 90 dB is dB.
See the Noise-Induced Hearing Loss guide for usage.
References
Section titled “References”- Ecma International. (2024). Psychoacoustic metrics for ITT equipment — Part 1: Prominent discrete tones (ECMA-418-1:2024 (3rd ed.)). The critical-band model and the TNR and PR procedures of the tone-prominence section. The linked PDF is the free download.
- Fastl, H., & Zwicker, E. (2007). Psychoacoustics: Facts and models (3rd ed.). Springer. https://doi.org/10.1007/978-3-540-68888-4The critical-band, masking and loudness psychoacoustics underneath the Zwicker model and the sharpness and roughness sensations.
- Fletcher, H., & Munson, W. A. (1933). Loudness, its definition, measurement and calculation. The Journal of the Acoustical Society of America, 5(2), 82-108. https://doi.org/10.1121/1.1915637The original equal-loudness measurements behind the contour concept of the first section.
- French, N. R., & Steinberg, J. C. (1947). Factors governing the intelligibility of speech sounds. The Journal of the Acoustical Society of America, 19(1), 90-119. https://doi.org/10.1121/1.1916407The articulation-band experiments behind the SII band-importance function.
- Houtgast, T., & Steeneken, H. J. M. (1985). A review of the MTF concept in room acoustics and its use for estimating speech intelligibility in auditoria. The Journal of the Acoustical Society of America, 77(3), 1069-1077. https://doi.org/10.1121/1.392224The modulation-transfer framework of the STI section.
- International Organization for Standardization. (2005). Acoustics — Reference zero for the calibration of audiometric equipment — Part 7: Reference threshold of hearing under free-field and diffuse-field listening conditions (ISO 389-7:2005). The Table 1 free-field and diffuse-field reference thresholds of the hearing-threshold section.
- International Organization for Standardization. (2013). Acoustics — Estimation of noise-induced hearing loss (ISO 1999:2013). The NIPTS model, its fractiles and the HTLAN compressed sum of the hearing-loss section.
- International Organization for Standardization. (2017). Acoustics — Statistical distribution of hearing thresholds related to age and gender (ISO 7029:2017). The age-dependent median shift and fractile spreads of the presbycusis model.
- International Organization for Standardization. (2023). Acoustics — Normal equal-loudness-level contours (ISO 226:2023). The Formula (1)/(2) contour model and the Table 1 parameters of the equal-loudness section.
- Passchier-Vermeer, W. (1974). Hearing loss due to continuous exposure to steady-state broad-band noise. The Journal of the Acoustical Society of America, 56(5), 1585-1593. https://doi.org/10.1121/1.1903482A field study of the exposure-response relations later codified in ISO 1999.