Perception and Hearing
Standards: ISO 226ECMA-418ISO 532DIN 45692IEC 60268ANSI S3.5Key references: Fletcher & Munson 1933Fastl & Zwicker 2007Houtgast & Steeneken 1985+5 more
This page collects the theory behind hearing and psychoacoustics: the equal-loudness contours, the Zwicker, Moore-Glasberg and Sottek loudness models, the sound-quality metrics tonality, roughness and sharpness, tone prominence, the speech metrics STI and SII, and the statistics of hearing thresholds and hearing loss. It is part of the theory reference.
Each section keeps the notation of the standard it describes, and five letters are reused between them. Where two collide the section says which is meant, but the table is here for anyone scrolling back to a formula:
| Symbol | Meaning | Section |
|---|---|---|
| Count of tone-band bins | Tone prominence (ECMA-418-1) | |
| Total loudness, in sone | Zwicker and advanced loudness (ISO 532) | |
| Noise-induced permanent threshold shift, in dB | Hearing loss (ISO 1999) | |
| Smoothed loudness state | Moore-Glasberg time-varying (ISO 532-3) | |
| Sharpness, in acum | Sharpness (DIN 45692) | |
| , threshold of hearing at , in dB | Equal-loudness contours (ISO 226) | |
| Tonality, in tu | Sottek Hearing Model (ECMA-418-2) | |
| Magnitude of the linear transfer function, normalised at 1 kHz | Equal-loudness contours (ISO 226) | |
| Mean level of the upper contiguous critical band | Tone prominence, PR (ECMA-418-1) | |
| Loudness level, in phon | Equal-loudness contours, Zwicker |
Equal-loudness contours (ISO 226:2023)
Section titled “Equal-loudness contours (ISO 226:2023)”A tone has a loudness level of phon when it is judged equally loud as a 1 kHz pure tone at dB SPL. ISO 226:2023 Formula (1) (clause 4.1, p. 2) gives the SPL of a pure tone at frequency that reaches loudness level :
Formula (2) (clause 4.2) inverts it, returning the loudness level of a tone at SPL :
The three parameters come from Table 1 (p. 4), tabulated at the 29 preferred third-octave frequencies of ISO 266 from 20 Hz to 12.5 kHz:
- : exponent for loudness perception at frequency ,
- : magnitude of the linear transfer function, normalized at 1 kHz ( at 1 kHz),
- : threshold of hearing at , in dB.
The standard specifies no interpolation between the tabulated frequencies. Formula (1) is specified for 20 phon to 90 phon between 20 Hz and 4 kHz, and only up to 80 phon between 5 kHz and 12.5 kHz; above 80 phon the contour therefore stops at 4 kHz. Values outside these limits from Formula (2) are extrapolations the standard labels as informative only.
See the Loudness guide for usage.
The ISO 226:2023 contours from Formula (1), 20 to 90 phon, with the hearing threshold.
Zwicker loudness (ISO 532-1)
Section titled “Zwicker loudness (ISO 532-1)”The ear analyzes sound in critical bands: frequency regions within which energy is summed before loudness is formed. The Bark scale maps frequency to critical-band rate , 0 to 24 Bark, and ISO 532-1:2017 samples the specific loudness at 0.1-Bark steps (240 values). The implementation is a clean-room port of the standard’s normative reference program (Annex A.4) and proceeds in stages:
-
One-third-octave levels: 28 bands, 25 Hz to 12.5 kHz (the Annex A filterbank at 48 kHz, Tables A.1/A.2). For time-varying sounds the squared band outputs are smoothed by three cascaded low-passes with ( capped at 1 kHz) and sampled every 2 ms.
-
Low-frequency grouping: the 11 bands up to 250 Hz receive the equal-loudness corrections of Table A.3 and are summed into the first three critical bands (25–80, 100–160, 200–250 Hz).
-
a0 transmission: the outer/middle-ear transfer correction of Table A.4 (plus the diffuse-field difference of Table A.5 when
field='diffuse') yields the critical-band levels . -
Core loudness: each of the 20 critical bands is transformed with the threshold-in-quiet levels of Table A.6 (after the bandwidth adaptation DCB of Table A.7):
(the reference program’s form of Zwicker’s loudness transformation; bands below threshold contribute zero).
-
Slopes: level-dependent upper masking slopes (steepness per specific-loudness range and critical band, Tables A.8/A.9) attach decaying flanks toward higher ; the total loudness is the area under the pattern:
For time-varying sounds a nonlinear temporal decay (time constants 5/15/75 ms, clause 6.3) and the duration-dependent weighting of the total loudness (3.5 ms and 70 ms low-passes weighted 0.47/0.53, clause 6.4) precede the 500 Hz loudness-vs-time output and the percentile values N5/N10 (clause 6.5).
Sone and phon are tied together by the 1 kHz anchor (1 sone = 40 phon; clause 5.6):
below 1 sone the reference program uses , floored at 3 phon.
See the Loudness guide for usage.
Specific loudness N′(z) over the Bark axis: energy spread over many critical bands sums to more sones than the same band level in a single band.
Advanced loudness models & sound quality
Section titled “Advanced loudness models & sound quality”ISO 532-1 is one of three loudness models; two newer families refine the auditory front-end and add the sound-quality metrics tonality and roughness.
All three are calibrated so that a 1 kHz tone at 40 dB SPL is 1 sone, so they agree at that anchor and diverge away from it — most visibly in level dependence, where Zwicker doubles the sone value every 10 dB and the Sottek model grows by about 1.65× — and in their treatment of narrowband against broadband signals. Sone values from different models are therefore not interchangeable: the model belongs beside every number, and two models must never be mixed in one comparison. The division of labour is what usually settles the choice. ISO 532-1 is the stationary workhorse that most product-noise regulations still name, and it is the loudness the sharpness definition below is built on. ISO 532-2 refines the auditory filters and adds binaural summation, still for stationary signals. ISO 532-3 carries that chain into the time domain for sounds up to a few seconds. ECMA-418-2 is the route to take when tonality and roughness have to come from the same front end as the loudness. When two models disagree, the useful question is which one the receiving standard names, not which is nearer the truth.
Moore-Glasberg loudness (ISO 532-2:2017, ISO 532-3:2023)
Section titled “Moore-Glasberg loudness (ISO 532-2:2017, ISO 532-3:2023)”Instead of Zwicker’s fixed critical bands, the Moore-Glasberg model forms a continuous excitation pattern on the ERB-number (“Cam”) scale using level-dependent rounded-exponential (roex) auditory filters. As a function of the normalized frequency deviation from a filter centred at , the filter weighting is
where the slope grows with the source level, broadening the lower skirt as level rises (ISO 532-2, Formulae 2–5); this reproduces the upward spread of masking. Passing the stimulus intensity through every filter gives the excitation , and a compressive law maps it to the specific loudness in sone/Cam (Formulae 7–9), of the mid-level form
with the calibration constant sone/Cam (ISO 532-2; in ISO 532-3). The total loudness is the area under the pattern,
and a binaural-inhibition stage (Formulae 10–13) combines the ears so a diotic sound is louder than the same sound at one ear. The 1 kHz / 40 dB SPL anchor gives exactly 1 sone.
ISO 532-3 makes this time-varying. A running spectrum from six parallel Hann-windowed FFTs (segment lengths 2–64 ms, each contributing its own frequency range, updated every ms) drives the same excitation and specific-loudness chain, integrated by two cascaded first-order smoothers with ,
using a fast time constant on the attack and a slower one on the release. This yields the short-term loudness (attack/release near 20–30 ms) and the long-term loudness (near 0.1–0.75 s); the peak long-term loudness predicts the loudness of sounds up to about 5 s.
See the Advanced Loudness guide for usage and the model-choice table.
Sottek Hearing Model (ECMA-418-2:2025)
Section titled “Sottek Hearing Model (ECMA-418-2:2025)”ECMA-418-2 builds all three of its metrics on one auditory front-end (Clause 5): an outer/middle-ear filter, a bank of 53 overlapping gammatone-like band-pass filters spaced on the Bark_HMS scale ( to ), half-wave rectification, and a short-block RMS per band and time block . A compressive nonlinearity (Formula 23) turns the band RMS into the specific basis loudness , whose calibration constant fixes a 1 kHz / 40 dB SPL tone at 1 sone_HMS. The loudness assembles the tonal and noise loudness (below) over bands and time (Formulae 113–117); it grows about per 10 dB, more slowly than Zwicker’s factor of 2, an intrinsic property of the Sottek summation.
The divergence stated above, drawn on the one signal all three models are calibrated on. They pass together through about 1 sone at 40 dB (the Sottek front end returns 0.9845 sone there) and then separate: Zwicker’s curve doubles every 10 dB, the Sottek curve rises by about 1.65× over the same step. That is a difference between auditory summations, not a calibration error, and it is why a sone value without its model attached cannot be compared with anything.
Tonality: autocorrelation of the band signal (ECMA-418-2)
Section titled “Tonality: autocorrelation of the band signal (ECMA-418-2)”A tonal component is periodic, so it survives in the autocorrelation function (ACF) of a band’s rectified signal while broadband noise decorrelates. For each band the unbiased ACF of the block is
A windowed spectral estimate of separates a tonal loudness from the noise loudness (Formulae 36–48). The specific tonality is the tonal loudness scaled by a smooth signal-to-noise gate (Formulae 49–51),
and the single value (tu_HMS) is the gated time-average of the per-block maximum over bands (Formulae 61–64). The constant fixes the 1 kHz / 40 dB tone at 1 tu_HMS, and the band of the ACF peak gives the tonal frequency .
Roughness: envelope modulation (ECMA-418-2)
Section titled “Roughness: envelope modulation (ECMA-418-2)”Roughness is the sensation of fast (roughly 20–300 Hz) amplitude modulation, strongest near 70 Hz. From each band’s envelope (Hilbert magnitude), a modulation spectrum is formed and weighted by a modulation-rate function peaking near 70 Hz and by the modulation depth; correlating the modulation across neighbouring bands and applying the specified temporal filtering yields the specific roughness and the time-dependent roughness
(Formulae 65–111). The single value is the 90th percentile of over time (Clause 7.1.10); the constant (Formula 104) calibrates the reference sound (a 1 kHz carrier 100 % amplitude-modulated at 70 Hz at 60 dB SPL) to 1 asper.
The modulation-rate weighting the formulae above apply, and the reason the range “roughly 20–300 Hz, strongest near 70 Hz” is a band-pass and not a threshold: the same 1 kHz carrier modulated slowly is heard as fluctuation strength, peaking near 4–6 Hz, and modulated fast is heard as roughness, peaking near 70 Hz. Between the two peaks the sensation changes name, not degree.
Sharpness (DIN 45692)
Section titled “Sharpness (DIN 45692)”Sharpness condenses the high-frequency emphasis of a sound into one number: the -weighted first moment of the ISO 532-1 stationary specific-loudness pattern (DIN 45692:2009, Equation 1):
evaluated on the same 240-bin, 0.1-Bark grid. The constant is not hard-coded but derived from the calibration requirement (clause 6): a critical-band-wide narrowband noise 920–1080 Hz at 60 dB SPL scores exactly 1 acum, and the derived lands inside the normative window (clause 5.2). The informative Annex B weightings are provided under the same 1-acum anchor: von Bismarck (knee at 15 Bark, ) and Aures (loudness-dependent, ). The Table A.2 narrow-band targets are reproduced within the clause 6 tolerance (5 % or 0.05 acum): 0.38 acum at 250 Hz, 1.00 at 1 kHz, 1.78 at 2.5 kHz, 2.82 at 4 kHz.
The three weightings of the formula above on one axis: DIN with its 15.8 Bark knee, von Bismarck with its 15 Bark knee, and the loudness-dependent Aures curve, which is why the choice of weighting changes a sharpness value only for sounds with energy above the knee (15 to 15.8 Bark, about 2.5 to 3 kHz) and leaves everything below it untouched.
What these three numbers are for. Each unit is anchored on a reference sound — 1 tu for a 1 kHz tone at 40 dB, 1 asper for a fully modulated 70 Hz-modulated tone, 1 acum for the 920–1080 Hz narrowband noise at 60 dB — precisely because no standard states an absolute criterion for any of them. There is no roughness or tonality that “fails”. They earn their keep as differences between design variants of the same product, measured the same way, with the same model. Orders of magnitude for reading a result: a fan or motor with an audible whine typically lands a few tenths of a tu, reported with its tonal frequency beside it, and a broadband sound with no periodic component sits near zero; a roughness of a few tenths of an asper is already an obtrusive beating or buzz, since 1 asper is a fully modulated tone at the worst modulation rate; sharpness near or below 1 acum is unremarkable and values above about 2 acum read as shrill. Differences of the order of 10 % in any of the three are around the smallest a listener reliably notices, so quoting more precision than that invites over-interpretation.
See the Sound Quality Metrics guide for usage.
Tone prominence: TNR and PR (ECMA-418-1)
Section titled “Tone prominence: TNR and PR (ECMA-418-1)”Both methods operate on a Hann-windowed, RMS-averaged power spectrum (clauses 11.1 / 12.1) and use the clause 10 critical-band model. The critical bandwidth centred on a tone at is (Formula 2):
Band edges are placed arithmetically for Hz (Formulae 4–5): , and geometrically above (Formulae 7–8): , .
TNR (clause 11). The tone band spans the spectral minima on both sides of the peak within 15 % of (clause 11.2). The tone power subtracts the straight line connecting the band-edge bins (Formula 9): over tone-band bins, . The masking-noise power is the remaining critical-band power rescaled to the full critical bandwidth (Formula 10): , and (Formula 11). The prominence criterion (Formulae 12–13) is
PR (clause 12) compares the level of the critical band centred on the tone, , with the mean power of the two contiguous critical bands , (edges from the fitted Formulae 21–22 with Tables 2–3): (Formula 23). For Hz the lower band is truncated at 20 Hz and its power rescaled to a 100 Hz bandwidth (Formula 24). The criterion (Formulae 25–26) is 9.0 dB at kHz, rising as below. Tones are assessed within the 89.1 Hz – 11.2 kHz range of interest (clauses 11.5 / 12.6).
The TNR criterion drawn rather than evaluated, over the 89.1 Hz – 11.2 kHz range of interest, with one assessed tone on it. Because the criterion is below 1 kHz and flat above, the same tone-to-noise ratio is judged against a different threshold at every frequency: the example tone clears its 13.0 dB threshold at 250 Hz by 2.1 dB, while a 10 dB tone would be prominent anywhere above 1 kHz and not prominent here.
See the Prominent Discrete Tones guide for usage.
Modulation transfer and STI (IEC 60268-16)
Section titled “Modulation transfer and STI (IEC 60268-16)”Speech intelligibility rides on the slow intensity modulations of the speech envelope. The clause numbers below are those of IEC 60268-16 Ed. 5 (2020). The modulation transfer function of a transmission channel is the ratio of received to emitted modulation depth of the octave-band intensity envelope at modulation frequency ; the full STI evaluates it at the 14 one-third-octave modulation frequencies 0.63–12.5 Hz in the seven octave bands 125 Hz – 8 kHz (A.2.2). From a measured impulse response the Schroeder closed form gives it directly (indirect method):
Steady background noise multiplies each band’s by the intensity ratio (the noise term):
and when absolute band levels are known the full correction adds the auditory masking intensity (from the next lower octave band, Table A.2) and the absolute reception threshold (Table A.3). Those two are absolute-level terms, so they only mean anything when the chain is calibrated in dB SPL at the listening position and the source is driven at a defined speech level. That splits the measurement in two: an uncalibrated impulse response gives , and with a separately measured background spectrum it gives the noise term as well, but the masking and reception-threshold terms are then left out rather than guessed; a calibrated measurement with the source radiating the standard speech spectrum at speaking height gives the full correction. It is also why a direct STIPA reading from a hand-held meter and an indirect STI from an impulse response of the same room can disagree when the source level was never set. Each corrected maps to an effective SNR, clipped to the ±15 dB range where intelligibility actually varies, then to a transmission index (A.5.4/A.5.5):
The band MTI is the mean TI over the modulation frequencies, and the STI weights the bands with the male factors , of Table A.1 (A.5.6):
truncated to 1.0.
The seven the weighted sum above consumes, for a hall with s and a 15 dB speech-to-noise ratio. Each bar is already the mean of 14 transmission indices, so this is two stages of averaging below the raw ; the bars sit within 0.06 of one another, which is the case in which the redundancy terms subtract almost nothing and the STI is close to the plain -weighted mean.
The end of the chain rather than its middle: what the Schroeder closed form does to the STI as reverberation grows, against the Annex F rating bands. The curve falls steeply through the range where a room is still usable and flattens once the modulation has already been destroyed, which is why halving a long reverberation time buys less intelligibility than halving a short one.
STIPA (Annex B) samples the same physics with just two modulation frequencies per band (Table B.1) on a test signal with source modulation index 0.55; the received depths are measured by sine/cosine correlation of the ~100 Hz low-passed intensity envelopes over an integer number of modulation periods:
Reading the index. STI is scaled 0 to 1 across the range over which
intelligibility actually varies, and the informative Annex F qualification bands
cut that range into 0.04-wide letters with edges at 0.36, 0.40 … 0.76, from U
below 0.36 to A+ at 0.76 and above; STIResult.rating returns the letter, and
Annex G Table G.1 gives an example use for each band — band G (0.48-0.52) is
labelled the target value for voice-alarm systems, which is where the familiar
“STI ≥ 0.5” requirement comes from, and H (0.44-0.48) the normal lower limit
for such systems. Two cautions come with the letter. Annexes F and G are
informative, and Edition 5’s own Scope says the document does not provide
criteria for certifying a transmission channel, so a project requirement belongs
in the contract as a numeric STI with the letter used to report it. And the
0.04 band width is itself set by the typical uncertainty of a direct
measurement, which for STIPA runs about 0.02 to 0.03 — wider than the 0.02
half-width of a band — so a value within about 0.02 of an edge may legitimately
be graded either side, and the STI is quoted beside the letter and averaged over
several positions rather than trusted at one.
See the Speech Transmission Index guide for usage, the Annex G table of typical uses and the position-averaging rule.
Speech Intelligibility Index (ANSI S3.5)
Section titled “Speech Intelligibility Index (ANSI S3.5)”Where the STI characterizes a transmission channel, the SII (ANSI S3.5-1997; the clause numbers below are that edition’s) predicts intelligibility from what the listener can actually hear: 18 one-third-octave bands 160 Hz – 8 kHz, each contributing its band importance (Table 3, , peaking near 2 kHz).
The importance function is where the speech cues are, and all four band procedures below agree about it: the same rise to a maximum near 2 kHz, the same total of 1.0 redistributed over wider or narrower bands — which is why the octave steps stand at 0.265 where the one-third-octave steps stand at 0.090. A band lost at 2 kHz costs several times what the same band costs at 160 Hz.
All inputs are equivalent spectrum levels (clauses 3.11/3.55), that is per-hertz quantities: a measured one-third-octave band level becomes or only after subtracting of that band’s bandwidth in hertz — about 23.6 dB in the 1 kHz band and over 30 dB at the top of the range — which is why the standard speech spectra sit near 25 dB at 1 kHz for a normal voice. Each of the three inputs comes from a defined measurement: the speech input is either one of the four built-in vocal efforts, defined at 1 m in front of the talker, or a measured speech spectrum referred to the listener’s position; the noise input is measured at that same listener position with the talker silent and over the same integration time; and the threshold input is the listener’s audiogram in dB HL at the same frequencies, zero for normal hearing. The check that catches the classic error is the anchor below: a normal-effort spectrum in quiet with normal hearing must score about 0.996, so an index that pins near 1.0 for an obviously noisy scene means band levels were passed where spectrum levels were expected.
Speech masks itself upward: each band’s masking spectrum (clause 5.4) accumulates the lower bands along slopes dB, and the disturbance is the larger of masking and hearing floor, (clause 5.6), with the reference internal noise spectrum plus the listener’s hearing-threshold shift (clauses 5.5/5.6). The band audibility clips the speech-to-disturbance margin into (clause 5.8), a level-distortion factor discounts overly loud presentation (clause 5.7), and the index sums (clause 6):
The same chain runs over the standard’s other three band tables, selected with method=: the 21 critical bands of Table 1, the 17 equally-contributing critical bands of Table 2 and the 6 octave bands of Table 4. Those three express the masking slope through the tabulated band width, , which is the same formula the above abbreviates for a one-third-octave band; the octave-band procedure omits the spread of masking altogether, since an octave band is already wider than the spread being modelled. The Table 3 standard speech spectra for the normal, raised, loud and shout vocal efforts are built in (25.01 / 33.86 / 42.16 / 51.31 dB at 1 kHz); in the level-distortion factor is always the normal-effort spectrum.
The four vocal efforts, in spectrum level rather than band level — which is why the normal-effort curve passes through 25.01 dB at 1 kHz rather than through the 60-odd dB a talker measures as a band level at a metre. Raising the effort does not lift the family uniformly: shouting adds 27.6 dB at 2.5 kHz, 26.3 dB at 1 kHz and −1.6 dB at 160 Hz, so the spectrum tilts as well as rises, and the extra effort is spent where the importance function above is largest.
Reference values: the normal-effort spectrum in quiet with normal hearing scores
SII = 0.996; the Table 3 masking reference values are reproduced within
dB of spectrum level; and the four vocal-effort spectra are verified
band by band against two independent public implementations, Google’s
speech_intelligibility_index and the R SII package (both in the references,
used for verification only).
Reading the index. The SII is an audibility fraction, not a score: it says what proportion of the importance-weighted speech cues reach the listener above masking and threshold. ANSI S3.5 maps it to a percentage of words or sentences correct only through transfer functions specific to the speech material and the listeners, so the same 0.5 supports high sentence scores and much lower nonsense-syllable scores, and an SII should never be quoted as “X % intelligible” without naming the material. As practical anchors: above about 0.75 is the design target where unfamiliar material must be understood, 0.45 to 0.75 is workable for familiar material in a known context, and below about 0.3 connected speech cannot be relied on. Differences of a few hundredths are inside the uncertainty of the input spectra and mean nothing. Which of the two speech metrics to reach for follows from what is being asked: the SII answers questions about a listener, because it is the only one of the two that takes an audiogram and a vocal effort as input; the STI answers questions about a room or a system, because it takes an impulse response and a background spectrum and says nothing about who is listening.
See the Speech Intelligibility guide for usage.
Hearing thresholds and presbycusis (ISO 389-7, ISO 7029)
Section titled “Hearing thresholds and presbycusis (ISO 389-7, ISO 7029)”ISO 389-7:2005 Table 1 fixes the reference threshold of hearing of otologically normal young adults: the free-field and diffuse-field SPL corresponding to 0 dB HL at the 11 audiometric frequencies 125 Hz – 8 kHz (22.1 dB at 125 Hz for both fields, 2.4/0.8 dB free/diffuse at 1 kHz, diverging at high frequency to 12.6 vs 6.8 dB at 8 kHz).
The two columns differ only in how the sound reaches the head, and the standard is exact about it (clause 1). Both apply to binaural listening by otologically normal 18-to-25-year-olds, and both are referred to the sound pressure level measured with the listener absent, at the point where the centre of the head would be. The free-field column applies to pure tones in a free progressive plane wave with the subject facing the source (frontal incidence); the diffuse-field column applies to one-third-octave bands of white or pink noise in a diffuse field — and up to 8 kHz to any narrower noise band as well. That difference in geometry is exactly why the columns coincide below 1 kHz and separate above it: at high frequency the frontal-incidence head and pinna gain exceeds the all-direction average, so less free-field level is needed for the same threshold. It also means these are calibration references for a sound-field audiometric setup (ISO 8253-2), not a curve to compare a room measurement against.
ISO 7029:2017 describes how that threshold shifts statistically with age: the median deviation from age 18 is (clause 4.2, Table 1)
and any fractile follows a two-sided Gaussian model (clause 4.4), , using the upper spread for (worse than median) and the lower spread otherwise, each a degree-5 polynomial in per sex and frequency (clause 4.3, Tables 2–5). At age 18 every deviation is zero by construction. The formulae are established to 80 years at and below 2 kHz and to 70 years above; beyond that the evaluation is an extrapolation. Reference values: at 60 years the medians evaluate to 7.85 dB (male, 1 kHz), 20.21 dB (male, 4 kHz) and 15.32 dB (female, 4 kHz), matching the Table 1 formula to .
See the Hearing Threshold guide for usage.
The ISO 7029 median age shift with its fractile band (left) and the ISO 389-7 free- and diffuse-field reference thresholds (right).
Noise-induced hearing loss (ISO 1999)
Section titled “Noise-induced hearing loss (ISO 1999)”ISO 1999:2013 predicts the permanent threshold shift a noise-exposed population accrues. The median noise-induced shift (NIPTS) for 10–40 years of exposure is (clause 6.3.1, Formula 2, Table 1):
quadratic in the excess over the frequency-dependent onset level (75 dB at 4 kHz, the most sensitive band, up to 93 dB at 500 Hz) and zero below it; under 10 years it scales as (Formula 3). Fractiles add the spread, with (clause 6.3.2, Formulae 4–7, Tables 2/3), clamped at zero. The library’s fractile counts the fraction of the population with the smaller shift, so fractile=0.9 is the most-susceptible decile (reliable range 0.05–0.95); ISO 1999’s own runs the other way, counting the percentage with worse hearing, so the same decile is there and in the Annex D column headings. The hearing threshold level associated with age and noise (HTLAN) combines NIPTS with the ISO 7029 age component at the same fractile through the compressed sum (clause 6.1, Formula 1):
The Annex D worked examples (Tables D.1–D.4; e.g. 100 dB / 40 yr at 3 kHz: 29/38/60 dB at the 0.10/0.50/0.90 fractiles) are reproduced exactly at the standard’s integer rounding, and the Formula 2 hand value at 4 kHz / 20 yr / 90 dB is dB.
The model as an audiogram: 40 years at an 8 h-normalised 95 dB(A). The notch
at 4 kHz is what makes noise-induced loss recognisable in a clinic, and it is
here only because is lowest (75 dB) in that band. Mind the fractile
direction the paragraph above states: the edge of the shaded band showing the
deeper shift is the library’s fractile=0.90, the most susceptible tenth
— which ISO 1999 and its Annex D column headings label . At 4 kHz
this case runs 19.5 / 26.0 / 36.0 dB at fractile 0.10 / 0.50 / 0.90.
See the Noise-Induced Hearing Loss guide for usage.
References
Section titled “References”- American National Standards Institute. (1997). Methods for calculation of the speech intelligibility index (ANSI S3.5-1997 (R2020)). Acoustical Society of America. The SII section: the clause 3.11/3.55 equivalent spectrum levels, the clause 5.4-5.8 masking and audibility chain, the clause 6 summation and the Tables 1-4 band procedures with their importance functions.
- Deutsches Institut für Normung. (2009). Measurement technique for the simulation of the auditory sensation of sharpness (DIN 45692:2009). The sharpness section: Equation (1), the clause 5.2 calibration window and the Table A.2 target values, with the Annex B weighting functions.
- Ecma International. (2024). Psychoacoustic metrics for ITT equipment — Part 1: Prominent discrete tones (ECMA-418-1:2024 (3rd ed.)). The critical-band model and the TNR and PR procedures of the tone-prominence section. The linked PDF is the free download.
- Ecma International. (2025). Psychoacoustic metrics for ITT equipment — Part 2: Models based on human perception (ECMA-418-2:2025 (4th ed.)). The Sottek Hearing Model: the Clause 5 auditory front end, the tonality of Clause 6, the roughness Formulae 113-117 and the sharpness input of Clause 7.1.10. The linked page carries the free PDF.
- Fastl, H., & Zwicker, E. (2007). Psychoacoustics: Facts and models (3rd ed.). Springer. https://doi.org/10.1007/978-3-540-68888-4The critical-band, masking and loudness psychoacoustics underneath the Zwicker model and the sharpness and roughness sensations.
- Fletcher, H., & Munson, W. A. (1933). Loudness, its definition, measurement and calculation. The Journal of the Acoustical Society of America, 5(2), 82-108. https://doi.org/10.1121/1.1915637The original equal-loudness measurements behind the contour concept of the first section.
- French, N. R., & Steinberg, J. C. (1947). Factors governing the intelligibility of speech sounds. The Journal of the Acoustical Society of America, 19(1), 90-119. https://doi.org/10.1121/1.1916407The articulation-band experiments behind the SII band-importance function.
- Google. (2019). speech_intelligibility_index. GitHub. One of the two independent public implementations the four vocal-effort spectra were cross-verified against band by band. Used for verification only.
- Houtgast, T., & Steeneken, H. J. M. (1985). A review of the MTF concept in room acoustics and its use for estimating speech intelligibility in auditoria. The Journal of the Acoustical Society of America, 77(3), 1069-1077. https://doi.org/10.1121/1.392224The modulation-transfer framework of the STI section.
- International Electrotechnical Commission. (2020). Sound system equipment — Part 16: Objective rating of speech intelligibility by speech transmission index (IEC 60268-16:2020 (Ed. 5)). The STI section: the A.2.2 modulation frequencies, the A.5.4/A.5.5 index mapping, the Table A.1 band weights, the Table A.2/A.3 masking and reception-threshold data, the Annex B STIPA method and the Annex F qualification bands.
- International Organization for Standardization. (2005). Acoustics — Reference zero for the calibration of audiometric equipment — Part 7: Reference threshold of hearing under free-field and diffuse-field listening conditions (ISO 389-7:2005). The Table 1 free-field and diffuse-field reference thresholds of the hearing-threshold section.
- International Organization for Standardization. (2013). Acoustics — Estimation of noise-induced hearing loss (ISO 1999:2013). The NIPTS model, its fractiles and the HTLAN compressed sum of the hearing-loss section.
- International Organization for Standardization. (2017). Acoustics — Methods for calculating loudness — Part 1: Zwicker method (ISO 532-1:2017). The Zwicker loudness section: the Annex A.4 reference program the implementation is a clean-room port of, its Tables A.1-A.7 and the clause 6.3 time-varying chain.
- International Organization for Standardization. (2017). Acoustics — Methods for calculating loudness — Part 2: Moore-Glasberg method (ISO 532-2:2017). The stationary Moore-Glasberg model of the advanced-loudness section: the roex auditory filters, the excitation pattern and the binaural summation.
- International Organization for Standardization. (2017). Acoustics — Statistical distribution of hearing thresholds related to age and gender (ISO 7029:2017). The age-dependent median shift and fractile spreads of the presbycusis model.
- International Organization for Standardization. (2023). Acoustics — Methods for calculating loudness — Part 3: Moore-Glasberg-Schlittenlacher method for time-varying signals (ISO 532-3:2023). The time-varying extension of Part 2 (Formulae 2-5 and 7-9), with the short-term and long-term loudness smoothers of the same section.
- International Organization for Standardization. (2023). Acoustics — Normal equal-loudness-level contours (ISO 226:2023). The Formula (1)/(2) contour model and the Table 1 parameters of the equal-loudness section.
- Passchier-Vermeer, W. (1974). Hearing loss due to continuous exposure to steady-state broad-band noise. The Journal of the Acoustical Society of America, 56(5), 1585-1593. https://doi.org/10.1121/1.1903482A field study of the exposure-response relations later codified in ISO 1999.
- Warnes, G. R. (2018). SII: Calculate ANSI S3.5-1997 Speech Intelligibility Index (version 1.0.3.1). CRAN archive. The second independent implementation used for the same band-by-band cross-verification of the vocal-effort spectra. Archived from CRAN in 2022; the last release is the one used.