Stone and Moore [(2014). J. Acoust. Soc Am. 135, 1967–77], showed that the introduction of explicit temporal-only modulations to a speech masker, that otherwise produced a near-constant envelope at the output of each auditory filter, rarely resulted in improved intelligibility, except at a very low modulation rate. This represents a failure in “dip-listening” or “glimpsing” [Cooke (2006). J. Acoust. Soc. Am. 119, 1562–1573], a facility where listeners are presumed to benefit from the temporarily improved signal-to-noise ratio during the masker dips. The dips of Stone and Moore only varied temporally, so Stone and Moore's method was used here to investigate the effect of maskers with both spectral and temporal dips, a pattern more representative of real-world maskers. For sinusoidally shaped modulations, intelligibility improved only at very low modulation rates, below 2 Hz temporally and 0.14 ripples/auditory filter spectrally. Square-wave modulation at a rate of 4 Hz resulted in improved intelligibility when only one cycle of spectral modulation was present across the audio bandwidth. Compared to the spectro-temporal extent of dips present during real-world noisy speech, dips generated by the reported modulation patterns were very large, further supporting the notion that dip-listening reflects a release from modulation masking and not energetic masking.