Skip to content

Overview

Reference material is generated from Python docstrings via mkdocstrings. The sections below mirror the main import paths used in applications and in the CLI.

Top-level package

Re-exports for common workflows (NoiseConfig, simulators, ParallelAdapter, lazy diagnostics/output symbols).

Top-level package for gwmock-noise.

ARMANoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate stationary detector noise with a fitted ARMA / state-space model.

denominator_coefficients property

Return a copy of the ARMA denominator coefficients a_1 .. a_p.

metadata property

Return metadata describing the fitted ARMA model.

model_psd_curve property

Return the fitted model PSD on the band-masked design grid.

numerator_coefficients property

Return a copy of the ARMA numerator coefficients b_0 .. b_q.

resume_metadata_nbytes property

Serialized size of the resume metadata in bytes.

state_nbytes property

Serialized size of the continuation state in bytes.

target_psd_curve property

Return the band-masked target PSD on the design grid.

__init__(*, psd_file=None, target_psd=None, target_frequencies=None, ar_order=DEFAULT_AR_ORDER, ma_order=DEFAULT_MA_ORDER, line_frequencies=None, line_widths=None, detect_lines=False, max_lines=DEFAULT_MAX_LINES, line_prominence=DEFAULT_LINE_PROMINENCE, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization=DEFAULT_REGULARIZATION)

Initialize the simulator and fit the ARMA model once.

Parameters:

Name Type Description Default
psd_file str | Path | None

Path or bundled preset name of a two-column PSD table.

None
target_psd ndarray | None

One-sided target PSD array, mutually exclusive with psd_file.

None
target_frequencies ndarray | None

Frequencies of target_psd in hertz.

None
ar_order int

Autoregressive order p of the denominator.

DEFAULT_AR_ORDER
ma_order int

Moving-average order q of the numerator.

DEFAULT_MA_ORDER
line_frequencies list[float] | None

Explicit line frequencies to place poles at.

None
line_widths list[float] | None

Widths matching line_frequencies; defaults to :data:DEFAULT_LINE_WIDTH_HZ.

None
detect_lines bool

Whether to detect prominent lines from the target.

False
max_lines int

Maximum number of detected lines.

DEFAULT_MAX_LINES
line_prominence float

Minimum peak-to-median ratio for detection.

DEFAULT_LINE_PROMINENCE
detectors list[str] | None

Detector names to generate.

None
sampling_frequency float

Sampling frequency in hertz.

4096.0
duration float

Nominal generation duration in seconds.

4.0
seed int | None

Base seed for the per-detector generators.

None
low_frequency_cutoff float

Lower edge of the target band in hertz.

2.0
high_frequency_cutoff float | None

Upper edge of the target band in hertz.

None
block_size int

Number of samples generated per filter call.

DEFAULT_BLOCK_SIZE
regularization float

Relative ridge added to the zero-lag autocovariance.

DEFAULT_REGULARIZATION

continuation_state()

Return the bounded state needed to keep a running stream going.

export_state()

Return a picklable snapshot sufficient to resume the stream.

from_component(component, config) classmethod

Construct an ARMA simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector ARMA noise with continuity across calls.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield ARMA chunks lazily while preserving filter state.

import_state(state)

Restore a snapshot produced by :meth:export_state.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same configuration.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator's settings.

reset()

Clear filter state and RNGs.

resume_metadata()

Return the metadata a stopped stream must persist to resume.

ARNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate stateful detector noise from an AR model fit to a target PSD.

autoregressive_coefficients property

Return a copy of the fitted a_1 .. a_p coefficients.

metadata property

Return metadata describing the fitted AR model.

model_psd_curve property

Return the fitted model PSD on the band-masked design grid.

resume_metadata_nbytes property

Serialized size of the resume metadata in bytes.

state_nbytes property

Serialized size of the continuation state in bytes.

target_psd_curve property

Return the band-masked target PSD on the design grid.

__init__(*, psd_file, order=DEFAULT_AR_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization=DEFAULT_REGULARIZATION)

Initialize the simulator and fit the AR model once.

Parameters:

Name Type Description Default
psd_file str | Path

Path or bundled preset name of a two-column PSD table.

required
order int

Autoregressive order p.

DEFAULT_AR_ORDER
detectors list[str] | None

Detector names to generate.

None
sampling_frequency float

Sampling frequency in hertz.

4096.0
duration float

Nominal generation duration in seconds.

4.0
seed int | None

Base seed for the per-detector generators.

None
low_frequency_cutoff float

Lower edge of the target band in hertz.

2.0
high_frequency_cutoff float | None

Upper edge of the target band in hertz.

None
block_size int

Number of samples generated per recursion call.

DEFAULT_BLOCK_SIZE
regularization float

Relative ridge added to the zero-lag autocovariance before the recursion; 0 disables it.

DEFAULT_REGULARIZATION

continuation_state()

Return the bounded state needed to keep a running stream going.

The state is the recursion memory of every detector; its size is proportional to the order and independent of the generated span.

export_state()

Return a picklable snapshot sufficient to resume the stream.

from_component(component, config) classmethod

Construct an AR-noise simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector AR noise with continuity across calls.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield AR-noise chunks lazily while preserving recursion state.

import_state(state)

Restore a snapshot produced by :meth:export_state.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same configuration.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator's settings.

reset()

Clear detector state and RNGs.

resume_metadata()

Return the metadata a stopped stream must persist to resume.

AddLines

Wrap a base simulator and add spectral lines to its output.

metadata property

Return additive-wrapper metadata.

__init__(base, lines)

Initialize the additive wrapper.

generate(duration, sampling_frequency, detectors, seed=None)

Generate base noise and add the configured spectral lines.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield additive spectral-line chunks lazily.

reset()

Reset the additive line state and any resettable base state.

BaseNoiseSimulator

Bases: ABC

Abstract base class for noise simulators.

This interface is the stable API through which the upstream gwmock package interacts with gwmock_noise. Implementations must override :meth:run.

run(config) abstractmethod

Run the noise simulation with the given configuration.

Parameters:

Name Type Description Default
config NoiseConfig

Validated noise simulation configuration.

required

Returns:

Type Description
SimulationResult

Result containing paths to generated outputs and the config used.

BlipGlitch dataclass

Bases: GlitchModel

Gaussian-windowed broadband burst.

__post_init__()

Validate blip-specific parameters.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate a Gaussian-windowed white-noise burst.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

ChannelDomain

Bases: StrEnum

The sample domain a channel array is expressed in.

ChannelKind

Bases: StrEnum

The physical role a channel plays in a joint realization.

ChannelMetadata dataclass

Typed, machine-checkable description of one channel.

Attributes:

Name Type Description
name str

Channel name, unique within one joint realization.

kind ChannelKind

Whether the channel is a gravitational-wave strain channel or an auxiliary/witness channel.

unit str

Physical unit of the channel's samples (e.g. "strain" for dimensionless strain, or a physical unit string such as "m/s^2").

domain ChannelDomain

Whether the channel's array is a time series or a frequency-domain series.

sampling_frequency float

Sampling frequency in hertz.

dtype str

The numpy dtype name (e.g. "float64") the channel array is generated in.

ColoredNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate colored detector noise from an input PSD.

metadata property

Return simulator metadata.

previous_strain property

Expose continuity buffers for protocol-compatible state inspection.

__init__(*, psd_file=None, psd_schedule=None, psd_array=None, frequencies=None, delta_frequency=None, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, window_duration=DEFAULT_WINDOW_DURATION)

Initialize the simulator.

export_state()

Return a picklable snapshot of the running crossfade stream.

The snapshot persists the bit-generator state and the chunk counter of the stitcher plus the bookkeeping needed to interpolate the PSD schedule on resume. It contains no cached strain: the previous window is regenerated from the bit-generator state.

from_component(component, config) classmethod

Construct a colored-noise simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector colored noise.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield colored-noise chunks lazily while preserving simulator state.

import_state(state)

Resume the crossfade stream from an :meth:export_state snapshot.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same settings.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator.

reset()

Clear continuity and RNG state.

CompositeNoiseSimulator

Additively combine multiple component simulators.

metadata property

Return metadata describing the composed run.

__init__(components, *, detectors, duration, sampling_frequency, seed)

Initialize the composed simulator.

generate(duration, sampling_frequency, detectors, seed=None)

Generate all component realizations and add them together.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield additive component chunks lazily.

ConfigurableNoiseSimulator

Bases: ABC

Abstract mixin for built-in simulators usable as composed components.

from_component(component, config) abstractmethod classmethod

Construct one simulator instance from a generic component config.

CorrelatedARNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate correlated detector noise with a truncated VMA representation.

metadata property

Return metadata describing the fitted VMA model.

__init__(*, psd_files, csd_files=None, order=DEFAULT_VMA_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization_epsilon=DEFAULT_REGULARIZATION_EPSILON)

Initialize the simulator and fit truncated VMA taps once.

from_component(component, config) classmethod

Construct a correlated-AR simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate correlated per-detector noise with stateful continuity.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield correlated AR chunks lazily while preserving innovation state.

reset()

Clear innovation state and RNG.

CorrelatedNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate correlated detector noise from PSD and CSD inputs.

metadata property

Return simulator metadata.

previous_strain property

Expose continuity buffers for protocol-compatible state inspection.

__init__(*, psd_files, csd_files=None, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, regularization_epsilon=1e-12, window_duration=DEFAULT_WINDOW_DURATION)

Initialize the simulator.

from_component(component, config) classmethod

Construct a correlated-noise simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate correlated per-detector noise.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield correlated-noise chunks lazily while preserving simulator state.

reset()

Clear continuity and RNG state.

DeepExtractorGlitch dataclass

Bases: GlitchModel

Real O3 glitch reconstructions colored against a target PSD.

Waveforms come from the DeepExtractor glitch-reconstruction dataset: whitened, amplitude-normalized 2-second time series sampled at 4096 Hz, covering seven Gravity Spy classes. Each drawn waveform is resampled to the simulation rate, colored against the configured PSD, and rescaled so its optimal SNR sqrt(4 df sum(|h(f)|^2 / S(f))) against that PSD matches the configured target. The dataset (~2.3 GB) is downloaded lazily on first use and cached by huggingface_hub; later runs reuse the cached files. On each run the Hub is contacted first to validate the cached files' ETags and download anything missing. If the Hub is unreachable, a warning notes that the ETag check was skipped and the cached files are used instead; if they are not cached, the first generate_waveform raises huggingface_hub's LocalEntryNotFoundError (a FileNotFoundError subclass). Set local_files_only=True to skip the network unconditionally and read straight from the cache.

revision pins the download to a specific dataset version — a git branch, tag, or commit SHA forwarded to hf_hub_download. Leave it None to track the repository default. Either way, the first download resolves to a concrete commit SHA (read back from the Hub cache path), which is what serialize records. Replaying a run through its metadata therefore fetches the exact commit that produced it, keeping glitch generation bit-reproducible for a fixed (version, config, seed) even as the upstream dataset moves.

rate accepts either a single number — the total Poisson rate shared by all configured classes, drawn uniformly — or a mapping from class name to per-class Poisson rate, in which case the total rate is their sum and each event's class is drawn proportionally to its rate.

snr accepts a number, a mapping from class name to number, a distribution, or a mapping from class name to distribution — freely mixed, so one class can be sampled while another stays fixed. A number calibrates every event of that class to exactly that loudness; a distribution draws a target per event from the model's own random stream, which is what a measured glitch population needs, since each Gravity Spy class is heavy-tailed and the tail indices differ by close to an order of magnitude between classes. See :mod:gwmock_noise.glitches.snr for the supported shapes.

What the target means is the caller's to decide. A tail index measured from a trigger generator's SNR — Omicron's, say — is defined against that pipeline's PSD over the band it searched, while this model calibrates against psd_file from low_frequency_cutoff upward. Configuring one from the other equates two SNRs defined against different noise curves over different bands. The sampler reproduces whatever distribution it is given and cannot tell whether that identification is the one you meant.

Resampling uses linear interpolation without an anti-aliasing filter, so sampling frequencies below 4096 Hz alias high-frequency content; the SNR calibration itself is unaffected because it is computed after resampling.

__post_init__()

Validate the configuration and preload the PSD table.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate one colored, SNR-calibrated DeepExtractor glitch.

resolve()

Pin the dataset to a concrete commit SHA, fetching only the samples file.

Reads the commit SHA huggingface_hub stamps into the cache path so the resolved revision is available for reproducibility metadata even when no waveform has been drawn yet — e.g. a batch that fires zero glitch events, where generate_waveform (and therefore _get_dataset) never runs. The samples file is downloaded once and reused by the later full load, so this adds no extra download on the generation path. Returns the resolved SHA, or None when the cache layout does not expose one (e.g. a test stub serving files from a flat directory).

serialize()

Return metadata-friendly model parameters.

DefaultNoiseSimulator

Bases: BaseNoiseSimulator

Default noise simulator implementation.

metadata property

Return metadata describing the current simulator state.

__init__(*, duration=4.0, sampling_frequency=4096.0, detectors=None, seed=None)

Initialize the simulator with protocol-compatible state.

generate(duration, sampling_frequency, detectors, seed=None)

Return Gaussian white-noise strain arrays.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield white-noise strain chunks lazily.

run(config)

Run the noise simulation with the given configuration.

DiagnosticResult dataclass

Outcome of a statistical diagnostic check.

EmpiricalSNRDistribution dataclass

Bases: SNRDistribution

Target SNR drawn with replacement from a table of observed SNRs.

For a caller holding the measurements rather than a fit to them: the drawn population is exactly the supplied one, tail included, with no assumption about its shape. Supply the values inline as samples or point file at a table on disk -- exactly one of the two.

Attributes:

Name Type Description
samples Sequence[float] | None

Observed SNRs given inline, or None when file is used.

file str | Path | None

Path to a table of observed SNRs, or None when samples is used. Read by :func:load_snr_samples.

distribution str

Discriminator naming this shape in a configuration.

__post_init__()

Validate the configuration and load the SNR table.

sample(rng)

Draw one target SNR uniformly from the table, with replacement.

serialize()

Return the mapping that reconstructs this distribution.

Echoes whichever of the two inputs was configured, so a file-backed table replays as a reference to the file rather than as a copy of its contents inlined into the metadata.

survival(threshold)

Return the fraction of the supplied table at or above threshold.

ExpectedCount dataclass

The expected injected count for one class in one interferometer, factorized.

The product is :attr:expected, and every factor is carried beside it. A campaign that finds the total surprising can then read off which factor is surprising -- a rate transplanted from another run, a livetime that is calendar time rather than observing time, a participation probability nobody meant to set -- instead of re-deriving the whole chain.

Attributes:

Name Type Description
glitch_class str

Name of the class within its population.

detector str

The interferometer this row is for.

rate_hz float

The class's Poisson rate. For a class whose interferometers glitch independently this is the rate each of them sees; for a network-coherent class it is the rate of the shared network process, which is not the same number unless every interferometer takes every event.

livetime_seconds float

The analysed time this count is over.

participation float

Probability that this interferometer receives a given event of the class. One for a class with no network coherence.

selection float

Fraction of the class's events that survive the SNR cut the count was asked for. One when no cut was given.

expected float

rate_hz * livetime_seconds * participation * selection.

network_events property

Expected number of network events of this class over the same livetime.

The same arithmetic with the participation factor divided out: the events the class's process produced, rather than the subset this interferometer received. For a class with no network coherence the two coincide, because every event it draws for an interferometer is that interferometer's.

FilterType

Bases: StrEnum

Types of segments that can be filtered out from GWOSC data.

FitError

Bases: ValueError

Raised when a model fit is numerically invalid or unstable.

GengliBlipGlitch dataclass

Bases: GlitchModel

File-backed gengli blip generator colored against a target PSD.

__post_init__()

Validate configured files and preload the population/PSD tables.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

from_population_file(population_file, **kwargs) classmethod

Construct a gengli glitch model from a population file path.

generate_waveform(sampling_frequency, rng=None)

Generate one colored gengli blip waveform.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

GlitchClass dataclass

One named class within a registered population.

The prose fields are part of the registered definition and therefore part of the digest, deliberately. A population whose numbers are unchanged but whose meaning has been rewritten is a different population, and a digest that did not move would say otherwise.

Attributes:

Name Type Description
name str

Stable identifier of the class within its population.

morphology str

Short slug naming the waveform family, e.g. "gaussian-windowed-broadband-burst".

description str

One sentence saying what the class represents and what it is drawn from.

model GlitchModel

The glitch model that generates it, already normalized.

__post_init__()

Validate the class definition.

Raises:

Type Description
ValueError

If the name, morphology or description is blank.

deserialize(payload) classmethod

Rebuild a class from :meth:serialize output.

Parameters:

Name Type Description Default
payload dict[str, Any]

The serialized class.

required

Returns:

Type Description
GlitchClass

The reconstructed class.

Raises:

Type Description
KeyError

If a required key is missing.

serialize()

Return the mapping that reconstructs this class.

GlitchDraw

Bases: NamedTuple

One drawn glitch waveform and the parameters that produced it.

What :meth:GlitchModel.draw returns so an injector can record what it injected, not merely that it injected something. Every field but the waveform is optional because not every model has the quantity: a parametric blip with no PSD has no SNR to report, and only the dataset-backed models draw a Gravity Spy class.

Attributes:

Name Type Description
waveform ndarray

The glitch strain, sampled at the requested rate. Its first sample lands at the injection time -- the waveform starts there, it does not peak there.

amplitude float | None

The amplitude multiplier drawn from the model's amplitude distribution, or None when the model does not draw one.

glitch_class str | None

The drawn morphological class (e.g. a Gravity Spy class name), or None for a model with a single morphology.

target_snr float | None

The optimal SNR the draw was calibrated to before the amplitude multiplier, or None when the model does not calibrate SNR. This is the configured target, not what the strain holds.

realized_snr float | None

The optimal SNR of the whole of waveform against the PSD it was colored with, or None for an uncolored model. Only the end of the generated data can leave the strain holding less than this; a waveform crossing a segment boundary has its remainder injected into the next segment.

How it relates to target_snr is model-specific, so do not assume one from the other. A model that calibrates against its own coloring PSD -- BlipGlitch, ScatteredLightGlitch, DeepExtractorGlitch -- realizes amplitude * target_snr. GengliBlipGlitch does not: its target is sampled from a population and imposed by gengli on the whitened waveform, while the realized figure is measured on the colored, amplitude-scaled result, so the two are independent numbers rather than one restated.

GlitchModel dataclass

Base dataclass for transient glitch generators.

detectors scopes the model to a subset of the run's interferometers. It accepts one name or a list of them, and defaults to None, which is every interferometer -- what a model without a selector has always meant. Scoping is what lets one configuration carry a 10 km model for one set of interferometers and a 15 km model for another, and per-interferometer rates for a network whose instruments are not equally glitchy (in O3, Fast_Scattering fired about 29 times more often in L1 than in H1).

A selector is only half of it. :func:validate_glitch_detector_coverage refuses a run whose models do not between them account for every interferometer, or whose models color one interferometer against two different PSDs -- see its docstring for why proceeding was the defect.

__post_init__()

Validate common glitch parameters.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate a single glitch waveform.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

GlitchNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate transient glitches as an additive standalone component.

from_component(component, config) classmethod

Construct a glitch-only component from one component definition.

GlitchPopulation dataclass

A named, versioned, digestible set of glitch classes.

The digest is only as portable as what goes into it. A model's psd_file is serialized as written, so a population built against /scratch/me/et.txt carries that path into its digest and no other machine can reproduce it. Name a bundled preset, or a path relative to the campaign's own tree, and the digest travels.

Attributes:

Name Type Description
name str

Stable registered identifier, e.g. "et-o3-anchored-v1".

description str

One paragraph saying what the population is for and what it is anchored to.

classes tuple[GlitchClass, ...]

The classes, in registration order.

references tuple[str, ...]

Published sources the registered quantities are anchored to, one citation per entry.

unanchored tuple[str, ...]

Quantities that could not be anchored to a published measurement, one per entry, each saying what it is and why. Carried in the population rather than only in its documentation, so a consumer reading a pinned population reads the caveats with it. An empty tuple is a claim that everything is anchored, so it is a claim to be able to defend.

schema_version str

Serialization schema, part of the digest.

__post_init__()

Validate the population definition.

Raises:

Type Description
ValueError

If the name or description is blank, the population has no classes, or two classes share a name.

build_injector(base, *, gps_start=0.0)

Wrap a base simulator so it also injects this population.

Parameters:

Name Type Description Default
base NoiseSimulator

The simulator producing the noise the glitches are added to.

required
gps_start float

GPS time of the first sample, stamped onto the truth catalogue.

0.0

Returns:

Type Description
InjectGlitches

The wrapped simulator.

canonical_json()

Return the exact bytes the digest is taken over.

Sorted keys, no insignificant whitespace, ASCII only and no NaN: a mapping that two processes can be relied on to render identically. Exposed rather than hidden so a campaign that disagrees about a digest can diff what was hashed instead of guessing.

Returns:

Type Description
str

The canonical serialization.

deserialize(payload) classmethod

Rebuild a population from :meth:serialize output.

Parameters:

Name Type Description Default
payload dict[str, Any]

The serialized population.

required

Returns:

Type Description
GlitchPopulation

The reconstructed population, which serializes and digests identically.

Raises:

Type Description
KeyError

If a required key is missing.

ValueError

If the payload declares a schema version this package cannot read.

digest()

Return the population's sha256:... digest.

What a campaign manifest pins. Every registered quantity, every prose field and the schema version go into it, so a population that has been edited cannot present itself as the one a finished run used.

Returns:

Type Description
str

The digest, algorithm-prefixed.

expected_counts(*, livetime_seconds, detectors, snr_threshold=None)

Decompose the expected injected count per class and interferometer.

Parameters:

Name Type Description Default
livetime_seconds float

Analysed time the count is over. This is observing time, not calendar time: the distinction is the commonest factor-of-1.3 error in a count derived from a published rate.

required
detectors Sequence[str]

The run's interferometers. Checked against the population's own scoping, so a decomposition cannot be computed for a network the population does not account for.

required
snr_threshold float | None

Optional recording or detection cut. When given, the selection factor is the fraction of each class's target-SNR distribution at or above it, read off the distribution in closed form.

None

Returns:

Type Description
list[ExpectedCount]

One row per (class, interferometer) pair, in registration order.

Raises:

Type Description
ValueError

If the livetime is not positive, no interferometers were given, or the population does not account for every interferometer.

model_for(glitch_class)

Return one class's glitch model by name.

Parameters:

Name Type Description Default
glitch_class str

The class name.

required

Returns:

Type Description
GlitchModel

The model that generates it.

Raises:

Type Description
KeyError

If the population has no such class.

models()

Return the population's glitch models in registration order.

realize(*, detectors, duration, sampling_frequency, seed, gps_start=0.0)

Generate this population's glitch strain on its own, with no noise under it.

This is the object paired arms share. A comparison whose arms differ in their signal content, their geometry, their detection threshold or their ranking statistic must not differ in their glitches, and the way to guarantee that is for every arm to add the same array rather than to re-run a generator and trust it to agree. The returned stamp carries the population digest and the seed, so an arm can record which realization it added -- and the digest is the field to pin, for the reason :class:GlitchRealization gives.

The realization is reproducible for a fixed (package version, population, seed) and does not depend on the base noise, on the other interferometers in the run, or on whether it was generated in one call or streamed -- see the module's tests, which assert each of those separately.

Parameters:

Name Type Description Default
detectors Sequence[str]

The interferometers to realize.

required
duration float

Length of the realization in seconds.

required
sampling_frequency float

Sampling frequency in Hz. Note that a class with a high_frequency_cutoff above the Nyquist frequency is refused rather than silently narrowed, so a population registered over a wide band needs a rate that can carry it.

required
seed int

The realization's seed.

required
gps_start float

GPS time of the first sample.

0.0

Returns:

Type Description
GlitchRealization

The glitch strain, its truth catalogue and the stamp identifying it.

Raises:

Type Description
ValueError

If the population does not account for every interferometer.

serialize()

Return the mapping that reconstructs this population.

GlitchRealization

Bases: NamedTuple

One seeded glitch realization and the stamp that identifies it.

Attributes:

Name Type Description
strain dict[str, ndarray]

The glitch contribution per interferometer, with nothing else in it. This is what paired arms add to their own base noise, and adding the same array to two arms is what makes their glitch content bit-identical.

events list[dict[str, Any]]

The per-event truth catalogue, in time order. See :data:gwmock_noise.simulators.glitches.GLITCH_CATALOGUE_COLUMNS.

stamp dict[str, Any]

What identifies this realization: the population's name and digest, the package version, and the generation parameters. Enough to reproduce strain and to prove another run used the same population.

The pinned value is digest, not the whole stamp. gwmock_noise_version is derived from the version-control description and therefore changes on every commit, including commits that touch nothing this population depends on; it is a provenance record of what produced a realization, and a manifest that compares it for equality would reject a rerun of the same population from a later commit. Pin population and digest, record the rest.

GwoscFilterConfig

Bases: BaseModel

Configuration for filtering GWOSC strain segments.

Attributes:

Name Type Description
filter_types list[FilterType]

Which categories of segments to exclude.

far_threshold float

Maximum false-alarm rate (events/year) for GW event filtering. Events with FAR above this threshold are excluded from filtering (treated as noise). Default 1.0 (one per year).

event_padding float

Padding in seconds to apply around each GW event GPS time when building vetosegments.

dq_flags list[str]

DQ flag basenames (without detector prefix) to query for data-quality vetosegments. E.g. ["CBC_CAT1", "CBC_CAT2"]. The detector prefix is prepended automatically.

exclude_hardware_injections bool

Whether to also exclude segments containing hardware injections.

has_dq_filter property

Return whether the data-quality filter is active.

has_gw_filter property

Return whether any GW signal filter is active.

include_marginal_events property

Return whether to include marginal (low-confidence) GW events.

GwoscNoiseConfig

Bases: BaseModel

Configuration for fetching real detector noise from GWOSC.

Attributes:

Name Type Description
detectors list[str]

List of detector prefixes (e.g. ["H1", "L1"]).

gps_start float

GPS start time of the requested data interval.

gps_end float

GPS end time of the requested data interval.

sample_rate float

Desired sampling rate in Hz. GWOSC typically provides 4096 Hz, with some event datasets at 16384 Hz.

filters GwoscFilterConfig

Filter configuration for excluding segments.

host str

GWOSC host URL.

cache_dir Path | None

Optional directory for file-level caching. When set, downloaded HDF5 frame files are saved locally and reused on subsequent requests for the same GPS interval. When None, data is downloaded on every call without caching.

duration property

Return the duration of the requested interval in seconds.

validate_gps_range()

Validate that gps_end is after gps_start.

GwoscNoiseFetcher

Fetch real detector noise data from GWOSC with optional filtering.

Downloads strain data from GWOSC and applies user-configured filters to exclude segments containing GW signals and data-quality issues.

When cache_dir is configured, HDF5 files are saved locally and reused on subsequent requests — avoiding repeated downloads for the same GPS interval.

Attributes:

Name Type Description
config

The GWOSC noise fetching configuration.

clean_segments property

Return the computed clean segments per detector.

Returns:

Type Description
dict[str, list[tuple[float, float]]]

A dictionary mapping each detector to a list of

dict[str, list[tuple[float, float]]]

(start, end) clean segment tuples.

__init__(config)

Initialize the fetcher.

Parameters:

Name Type Description Default
config GwoscNoiseConfig

Configuration specifying detectors, GPS range, sample rate, filtering options, and optional cache directory.

required

check_availability(detectors=None)

Probe GWOSC for per-detector strain-data availability.

Unlike :attr:clean_segments, which only reflects vetoes (GW events and data-quality flags), this checks whether the strain data itself has been published and is downloadable for the full configured GPS interval. A detector can have a fully "clean" interval yet have no published strain — e.g. the data has not been released yet for the observing run — in which case a fetch would fail. Use this for a pre-flight check before fetching.

Parameters:

Name Type Description Default
detectors list[str] | None

Detectors to probe. Defaults to all configured detectors when None.

None

Returns:

Type Description
dict[str, bool]

A dictionary mapping each requested detector to True if

dict[str, bool]

GWOSC has data URLs covering the interval, False otherwise.

fetch_clean(detectors=None)

Fetch clean noise segments.

Clean segments are computed by excluding GW events and data-quality issues according to the filter configuration.

Parameters:

Name Type Description Default
detectors list[str] | None

Detectors to fetch. Defaults to all configured detectors when None.

None

Returns:

Type Description
dict[str, list[TimeSeries]]

A dictionary mapping each requested detector to a list of

dict[str, list[TimeSeries]]

gwpy.TimeSeries, one per clean segment.

Raises:

Type Description
ValueError

If no data is available for any requested detector or no clean segments are found.

fetch_raw(detectors=None)

Fetch raw strain data without filtering.

Parameters:

Name Type Description Default
detectors list[str] | None

Detectors to fetch. Defaults to all configured detectors when None.

None

Returns:

Type Description
dict[str, TimeSeries]

A dictionary mapping each requested detector to a full-interval

dict[str, TimeSeries]

gwpy.TimeSeries.

Raises:

Type Description
ValueError

If no data is available for any requested detector.

GwoscNoiseSimulator

Real-noise simulator that fetches strain data from GWOSC.

Implements the NoiseSimulator protocol so it can be used interchangeably with synthetic simulators like ColoredNoiseSimulator and CorrelatedNoiseSimulator.

Fetches real detector strain from the Gravitational-Wave Open Science Centre and applies user-configured filters to exclude segments containing GW signals or data-quality issues.

When cache_dir is set in the config, downloaded HDF5 files are saved locally and reused — avoiding repeated downloads for the same GPS interval.

Attributes:

Name Type Description
duration float

Duration of the configured GPS interval (seconds).

sampling_frequency float

Sampling frequency in Hz.

detectors list[str]

List of detector prefixes.

seed None

Always None (real noise has no random seed).

config

The underlying GWOSC configuration.

detectors property

Return the list of detector prefixes.

duration property

Return the total duration of the configured GPS interval.

metadata property

Return metadata describing the simulator and its configuration.

Returns:

Type Description
dict[str, Any]

A dictionary with the simulator implementation name,

dict[str, Any]

GPS range, sample rate, detectors, filter configuration,

dict[str, Any]

and cache status.

sampling_frequency property

Return the sampling frequency in Hz.

seed property

Return None — real noise has no controllable random seed.

__init__(config)

Initialize the real-noise simulator.

Parameters:

Name Type Description Default
config GwoscNoiseConfig

Configuration specifying GPS range, detectors, sample rate, filtering options, and optional cache directory.

required

check_availability(detectors=None)

Probe GWOSC for per-detector strain-data availability.

Convenience passthrough to :meth:GwoscNoiseFetcher.check_availability. Use this as a pre-flight check: a detector may have a fully clean interval (no vetoes) yet have no published strain data, in which case :meth:generate would fail with a clear error.

Parameters:

Name Type Description Default
detectors list[str] | None

Detectors to probe. Defaults to all configured detectors when None.

None

Returns:

Type Description
dict[str, bool]

A dictionary mapping each requested detector to True if

dict[str, bool]

GWOSC has data covering the interval, False otherwise.

generate(duration, sampling_frequency, detectors, seed=None)

Fetch clean noise and return the longest contiguous segment per detector.

Real GWOSC noise is fragmented by the excision of GW events and data-quality vetoes, so the clean data generally consists of several non-contiguous spans. To honour the NoiseSimulator contract — a single, gap-free array per detector — this returns the longest contiguous clean segment for each detector.

Concatenating the fragments instead (the historical behaviour) splices non-adjacent samples together, creating discontinuities that show up as sharp transients after whitening. Use :meth:generate_segments to obtain every clean segment separately when the full clean data is required.

Parameters:

Name Type Description Default
duration float

Requested duration (ignored; the longest available contiguous clean segment determines the output length).

required
sampling_frequency float

Requested sampling frequency (must match the configured sample_rate).

required
detectors list[str]

Requested detector list (must be a subset of the configured detectors).

required
seed int | None

Ignored for real noise.

None

Returns:

Type Description
dict[str, ndarray]

A dictionary mapping each detector to a 1-D numpy array

dict[str, ndarray]

holding its longest contiguous clean segment.

Raises:

Type Description
ValueError

If sampling_frequency does not match the configured value, if detectors are not a subset, or if a detector has no clean data.

generate_segments(duration, sampling_frequency, detectors, seed=None)

Fetch clean noise and return per-detector lists of contiguous segments.

Unlike :meth:generate, this preserves the gap structure of the real data: each clean segment — the spans left after excluding GW events and data-quality vetoes — is returned as its own array, in GPS-time order. Adjacent segments are not contiguous in time, so callers must treat each segment independently; concatenating them would splice non-adjacent samples together and introduce discontinuities at the joins (which manifest as sharp transients after whitening).

Parameters:

Name Type Description Default
duration float

Requested duration (ignored; the GPS interval from the config determines the available data).

required
sampling_frequency float

Requested sampling frequency (must match the configured sample_rate).

required
detectors list[str]

Requested detector list (must be a subset of the configured detectors).

required
seed int | None

Ignored for real noise.

None

Returns:

Type Description
dict[str, list[ndarray]]

A dictionary mapping each detector to a list of 1-D numpy

dict[str, list[ndarray]]

arrays, one per clean segment, ordered by GPS time.

Raises:

Type Description
ValueError

If sampling_frequency does not match the configured value, if detectors are not a subset, or if a detector has no clean data.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield clean-noise chunks lazily.

Fetches the longest contiguous clean segment once (see :meth:generate) and yields it in chunks of chunk_duration seconds. Chunking the longest contiguous span keeps every chunk free of the discontinuities that splicing fragments would introduce.

Parameters:

Name Type Description Default
chunk_duration float

Duration of each yielded chunk in seconds.

required
sampling_frequency float

Requested sampling frequency (must match the configured sample_rate).

required
detectors list[str]

Requested detector list (must be a subset of the configured detectors).

required
seed int | None

Ignored for real noise.

None

Yields:

Type Description
dict[str, ndarray]

Per-detector strain arrays for each chunk.

GwoscSegmentFilter

Compute clean analysis segments from GWOSC metadata.

Queries GWTC event catalogs and data-quality flags to build vetosegments (time windows to exclude) and returns the remaining clean intervals suitable for noise analysis.

Attributes:

Name Type Description
config

The filter configuration.

__init__(config)

Initialize the segment filter.

Parameters:

Name Type Description Default
config GwoscFilterConfig

Filter configuration specifying which segment categories to exclude and their parameters.

required

compute_clean_segments(gps_start, gps_end, detectors)

Compute clean (analysis-ready) segments for each detector.

Clean segments are the requested GPS range minus the union of: - GW event vetosegments (if any GW filter is active) - DQ vetosegments (if the data-quality filter is active) - Hardware injection segments (if enabled)

Parameters:

Name Type Description Default
gps_start float

GPS start of the requested interval.

required
gps_end float

GPS end of the requested interval.

required
detectors list[str]

List of detector prefixes.

required

Returns:

Type Description
dict[str, SegmentList]

A dictionary mapping each detector to a list of clean

dict[str, SegmentList]

(start, end) segments.

get_dq_vetosegments(gps_start, gps_end, detector)

Query data-quality flags for detector and build vetosegments.

Uses gwosc.timeline.get_segments() to fetch pre-computed DQ veto segments for each flag in dq_flags.

Parameters:

Name Type Description Default
gps_start float

GPS start of the query interval.

required
gps_end float

GPS end of the query interval.

required
detector str

Detector prefix (e.g. "H1").

required

Returns:

Type Description
SegmentList

A list of (start, end) vetosegment tuples covering

SegmentList

time windows with data-quality issues.

get_gw_vetosegments(gps_start, gps_end)

Query GWTC events in the GPS range and build vetosegments.

For HIGH_CONFIDENCE_GW, only events with FAR <= far_threshold are included. For ALL_GW_SIGNALS, all events in the range are included.

Parameters:

Name Type Description Default
gps_start float

GPS start of the query interval.

required
gps_end float

GPS end of the query interval.

required

Returns:

Type Description
SegmentList

A list of (start, end) vetosegment tuples, each centred

SegmentList

on a GW event GPS time with event_padding on both sides.

get_hardware_injection_segments(gps_start, gps_end)

Query hardware injection events and build vetosegments.

Parameters:

Name Type Description Default
gps_start float

GPS start of the query interval.

required
gps_end float

GPS end of the query interval.

required

Returns:

Type Description
SegmentList

A list of (start, end) vetosegment tuples around

SegmentList

hardware injection times.

InjectGlitches

Wrap a base simulator and inject transient glitches additively.

A glitch model with no network coherence runs an independent Poisson process per detector, so every detector receives its own event times and waveform realizations, and GlitchModel.rate is the event rate seen by each individual detector. Every (model, detector) pair draws from its own random generator, derived from the top-level seed and the detector name via an independent SeedSequence stream.

A model carrying a :class:~gwmock_noise.glitches.models.NetworkCoherence runs one Poisson process for the whole network and draws one waveform per event, offering it to every detector it applies to. GlitchModel.rate is then the network event rate, not the per-detector one, and each detector sees it times the coherence's participation probability. Arrival times and waveforms come from a stream keyed by the model alone; each detector's participation and amplitude ratio come from a stream keyed by its own name and the event's ordinal.

Either way, a detector's realization is reproducible and independent of which other detectors are present or the order in which they are requested -- which is what lets two runs over different networks be compared on the channels they share.

A glitch whose waveform extends past the end of a chunk has its remainder carried into the next chunk and replayed in event order, so the injected glitch series is identical, sample for sample, to a single generate call of the same total duration (the per-sample addition order is preserved exactly, not merely up to rounding). A tail is dropped where the data it would continue into does not exist: past the last generated chunk, and past a segment whose successor the caller places at a discontinuous epoch.

Every event that lands in the strain is recorded as it fires, not merely tallied: see :attr:glitch_events for the cumulative truth catalogue and :attr:segment_glitch_events for the events the most recent generate call injected. A row is written where a glitch starts, so a waveform straddling a chunk boundary appears once, in the chunk that holds its first sample, at its true time -- the carried-over tail adds samples to the next chunk but no second row. GLITCH_CATALOGUE_COLUMNS documents the columns and GLITCH_CATALOGUE_TIME_CONVENTION which instant each time column names.

gps_start is the GPS time of the first sample of the next segment, and it belongs to the caller. It advances by each generated segment's duration, so a stream's catalogue reads in real GPS time on its own; a caller that knows better -- a writer placing non-contiguous segments, or one driving the stream batch by batch -- assigns it before each generate call and that assignment is honoured, including on a call that also carries a seed. Only reset, the explicit start-over, rewinds it to the epoch the wrapper was constructed with.

glitch_events property

Return every glitch injected since the last (re)seed, in time order.

Copies, so a consumer cannot edit the run's own record of itself.

metadata property

Return additive-wrapper metadata.

segment_glitch_events property

Return the glitches the most recent generate call injected.

What a per-batch consumer wants: the rows belonging to the chunk it is about to write, rather than the whole stream's catalogue re-filtered by time. Empty before the first generate call, and after a call that fired no events.

__init__(base, glitch_models, gps_start=0.0)

Initialize the additive glitch wrapper.

generate(duration, sampling_frequency, detectors, seed=None)

Generate base noise and inject the configured glitches.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield glitch-injected chunks lazily.

reset()

Reset the additive wrapper and any resettable base state.

Rewinds the epoch to the one the wrapper was constructed with, which _initialize_process does not: this is the caller saying "start over", where a seed passed to generate is the caller saying "this segment, whose epoch I have just told you, begins a new realization". The two need different answers, and conflating them is what put every streamed segment on the wrong epoch.

JointCovariance dataclass

A one-sided cross-spectral density matrix over a set of channels.

The convention matches :func:numpy.fft.rfftfreq: frequencies is the one-sided, non-negative frequency grid, and matrix[k] is the Hermitian (n_channels, n_channels) one-sided cross-spectral density at frequencies[k], in units of channel_unit**2 / Hz. The diagonal entry matrix[k, i, i] is the one-sided PSD of channel channel_order[i] at that frequency.

Attributes:

Name Type Description
frequencies ndarray

One-sided frequency grid in hertz, shape (n_freq,).

matrix ndarray

Cross-spectral density matrices, shape (n_freq, n, n), complex-valued and Hermitian at every frequency.

channel_order list[str]

Channel names, in the order matrix's last two axes index them.

JointDummyCorrelatedSimulator

Public dummy backend producing simultaneous, correlated strain and witness channels.

Every channel is temporally white with unit per-sample variance and coupling correlation with every other channel; there is no spectral shaping and no physical model. The one-sided cross-spectral density is therefore flat and known in closed form: for per-sample covariance Sigma at sampling frequency fs, the one-sided PSD/CSD is the frequency-independent matrix 2 * Sigma / fs (the same covariance-to-PSD scale convention already used by :class:~gwmock_noise.simulators.multichannel.MultichannelNoiseSimulator, inverted).

channel_metadata property

Return typed metadata for every strain and witness channel.

metadata property

Return simulator metadata.

__init__(*, detectors, witnesses, sampling_frequency=DEFAULT_SAMPLING_FREQUENCY, duration=DEFAULT_DURATION, seed=None, coupling=DEFAULT_COUPLING, strain_unit='strain', witness_unit='counts')

Initialize the dummy backend.

Parameters:

Name Type Description Default
detectors list[str]

Strain channel names. Fixed for the lifetime of this instance; later calls may reorder but not change this set.

required
witnesses list[str]

Witness channel names, disjoint from detectors. Fixed for the lifetime of this instance in the same way.

required
sampling_frequency float

Sampling frequency in hertz.

DEFAULT_SAMPLING_FREQUENCY
duration float

Nominal generation duration in seconds.

DEFAULT_DURATION
seed int | None

Seed for the shared innovation generator.

None
coupling float

Equicorrelation coefficient applied between every pair of channels (strain-strain, witness-witness and strain-witness alike). Must satisfy 0.0 <= coupling < 1.0.

DEFAULT_COUPLING
strain_unit str

Unit string recorded for strain channels.

'strain'
witness_unit str

Unit string recorded for witness channels.

'counts'

Raises:

Type Description
ValueError

If detectors/witnesses are empty, contain duplicates, overlap each other, or if coupling is out of range.

covariance()

Return this backend's exact, closed-form one-sided cross-spectral density.

Every channel is temporally white, so the returned matrix is frequency-independent: 2 * Sigma / sampling_frequency at every frequency, where Sigma is the per-sample covariance matrix drawn from at construction. This is an exact analytic value, not a statistical estimate off a realization.

generate_joint(duration, sampling_frequency, detectors, witnesses, seed=None)

Generate one simultaneous, correlated strain+witness realization.

generate_joint_stream(chunk_duration, sampling_frequency, detectors, witnesses, seed=None)

Yield simultaneous strain+witness chunks lazily, preserving RNG state.

JointRealization dataclass

One simultaneous draw of strain and witness channels.

Attributes:

Name Type Description
strain dict[str, ndarray]

Per-detector strain arrays, one entry per requested detector.

witness dict[str, ndarray]

Per-channel witness arrays, one entry per requested witness channel.

channel_metadata dict[str, ChannelMetadata]

Typed metadata for every channel in strain and witness, keyed by the same channel names.

provenance dict[str, Any]

Implementation-defined provenance describing how this realization was produced (backend identity, package version, seed, RNG algorithm, and any upstream configuration actually used).

JointStrainWitnessSimulator

Bases: Protocol

Structural interface for simulators producing joint strain+witness output.

This is an optional extension of the legacy NoiseSimulator contract: implementing it changes nothing about NoiseSimulator itself, and a NoiseSimulator-only consumer is unaffected by its existence. A class may satisfy both protocols at once.

metadata property

Return simulator metadata (same role as NoiseSimulator.metadata).

covariance()

Return the joint one-sided cross-spectral density over all channels.

generate_joint(duration, sampling_frequency, detectors, witnesses, seed=None)

Generate one simultaneous strain+witness realization.

generate_joint_stream(chunk_duration, sampling_frequency, detectors, witnesses, seed=None)

Yield simultaneous strain+witness chunks lazily.

Continuity contract: consecutive chunks from this iterator must equal the same realization a caller would obtain from one seeded :meth:generate_joint call spanning the combined duration -- the same contract NoiseSimulator.generate_stream makes for strain-only output.

LogNormalAmplitudeDistribution dataclass

Log-normal amplitude sampler parameterized by linear mean and std.

__post_init__()

Validate the configured distribution parameters.

sample(rng)

Draw one amplitude sample.

MultichannelNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate multichannel noise from a tabulated PSD/CSD matrix.

ar_coefficients property

Return a copy of the fitted coefficient matrices A_1 .. A_p.

design_frequencies property

Return a copy of the design-grid frequencies in hertz.

innovation_covariance property

Return a copy of the fitted innovation covariance.

metadata property

Return metadata describing the fitted matrix spectral factor.

model_spectral_matrices property

Return a copy of the fitted cross-spectral matrices on the design grid.

model_spectral_matrix_curve property

Return the band-masked fitted cross-spectral matrices on the design grid.

resume_metadata_nbytes property

Serialized size of the resume metadata in bytes.

state_nbytes property

Serialized size of the continuation state in bytes.

target_spectral_matrices property

Return a copy of the target cross-spectral matrices on the design grid.

The matrices are in the covariance scale the recursion uses (one_sided_psd * sampling_frequency / 2) and carry the band mask and edge taper the fit was built from.

target_spectral_matrix_curve property

Return the band-masked target cross-spectral matrices on the design grid.

__init__(*, psd_files=None, csd_files=None, target_matrices=None, target_frequencies=None, order=DEFAULT_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization_epsilon=DEFAULT_REGULARIZATION_EPSILON)

Initialize the simulator and fit the matrix spectral factor once.

The cross-spectral matrix may be given either as files (psd_files plus csd_files, absent off-diagonal pairs mean zero coherence) or as an in-memory target_matrices array on target_frequencies, which lets an analytic pair bypass the tabulated-curve interpolation.

Parameters:

Name Type Description Default
psd_files dict[str, str | Path] | None

Per-detector PSD table paths, keyed by detector.

None
csd_files dict[str, str | Path] | dict[tuple[str, str], str | Path] | None

Cross-spectral table paths, keyed by detector pair or by "DET1-DET2" strings.

None
target_matrices ndarray | None

One-sided cross-spectral matrices of shape (n_frequencies, n_channels, n_channels), mutually exclusive with the file inputs.

None
target_frequencies ndarray | None

Frequencies of target_matrices in hertz.

None
order int

Order p of the autoregressive factor.

DEFAULT_ORDER
detectors list[str] | None

Detector names to generate.

None
sampling_frequency float

Sampling frequency in hertz.

4096.0
duration float

Nominal generation duration in seconds.

4.0
seed int | None

Seed for the shared innovation generator.

None
low_frequency_cutoff float

Lower edge of the fit band in hertz.

2.0
high_frequency_cutoff float | None

Upper edge of the fit band in hertz.

None
block_size int

Number of samples generated per recursion call.

DEFAULT_BLOCK_SIZE
regularization_epsilon float

Relative ridge, as a fraction of the largest in-band diagonal entry, added as a zero-lag (white) floor to the autocovariance when the in-band target is not positive definite. Zero forbids the ridge and makes such a target a fit failure.

DEFAULT_REGULARIZATION_EPSILON

continuation_state()

Return the bounded state needed to keep a running stream going.

The state is the recursion history plus the pending output; its size is proportional to the order and independent of the generated span.

export_state()

Return a picklable snapshot sufficient to resume the stream.

from_component(component, config) classmethod

Construct a multichannel simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector multichannel noise with continuity across calls.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield multichannel chunks lazily while preserving recursion history.

import_state(state)

Restore a snapshot produced by :meth:export_state.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same configuration.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator's settings.

reset()

Clear the recursion history, pending output and generator.

resume_metadata()

Return the metadata a stopped stream must persist to resume.

NetworkCoherence dataclass

Share one glitch model's events across the interferometers it applies to.

Without it, every glitch model runs an independent Poisson process and an independent random stream per interferometer, so two interferometers never carry the same transient. That is the right default, and it is what the one measurement of the question says for widely separated sites: over O1 and O2 the LIGO blip population produced no coincidences inside the ±15 ms window an astrophysical signal can occupy, and the coincidences found in wider windows matched the accidental expectation.

It is not what a network of co-located interferometers is expected to do. Three interferometers sharing one site, one vacuum system and one seismic environment see a common environmental transient in all three at once, and a coincidence veto -- or a null stream -- assumes exactly the incoherence that such an event breaks. A model carrying this runs one Poisson process for the whole network, draws one waveform per event, and offers that waveform to each interferometer it applies to.

Participation is decided per interferometer, independently, and is never conditioned on how many others took the event. Conditioning -- "keep only events that land in at least two interferometers" -- is the obvious way to write a coincident population, and it is the wrong one here: it makes what one interferometer's strain contains depend on which other interferometers the run happens to include, so the same channel stops being reproducible between a three-interferometer run and a two-interferometer one. Paired-geometry comparisons need that reproducibility, so an event that lands nowhere simply lands nowhere, and the multiplicity comes out Binomial rather than being imposed.

Attributes:

Name Type Description
participation_probability float

Probability that any one interferometer the model applies to receives a given network event. 1.0 -- the default -- puts every event in every one of them, which is the strongest coherence the model can express.

amplitude_ratio_std float

Standard deviation of a per-interferometer lognormal amplitude ratio with linear mean 1.0, applied to the shared waveform. 0.0 -- the default -- injects the identical strain into every participating interferometer.

__post_init__()

Validate the configured coherence parameters.

Raises:

Type Description
ValueError

If the participation probability is outside [0, 1] or the amplitude-ratio spread is negative.

amplitude_ratio(rng)

Draw one interferometer's amplitude ratio for a shared waveform.

Parameters:

Name Type Description Default
rng Generator

The generator for this (event, interferometer) pair.

required

Returns:

Type Description
float

The multiplier to apply to the shared waveform.

serialize()

Return the mapping that reconstructs this coherence specification.

NoiseComponentConfig

Bases: BaseModel

One configurable noise component in a composed simulation.

normalize_component_definition(value) classmethod

Accept string, flat mapping, or explicit {simulator, options} input.

validate_component()

Validate the normalized component entry.

NoiseConfig

Bases: BaseModel

Generic configuration for composed detector-noise simulations.

NoiseSimulator

Bases: Protocol

Structural interface for downstream noise simulators.

metadata property

Return simulator metadata.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector strain arrays.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield per-detector strain chunks lazily.

OutputConfig

Bases: BaseModel

Configuration for simulation output.

OverlapSaveFirSimulator

Bases: ConfigurableNoiseSimulator

Generate coloured detector noise with a minimum-phase overlap-save FIR.

filter_taps property

Return a copy of the designed colouring filter taps.

metadata property

Return simulator metadata.

resume_metadata_nbytes property

Serialized size of the resume metadata in bytes.

state_nbytes property

Serialized size of the continuation state in bytes.

target_psd property

Return a copy of the band-masked target PSD the filter was designed from.

__init__(*, psd_file=None, target_psd=None, target_frequencies=None, filter_length=DEFAULT_FILTER_LENGTH, minimum_phase=True, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE)

Initialize the simulator and design the colouring filter once.

The target may be given either as a file (psd_file) or directly as a one-sided array (target_psd), so an analytic target can bypass the tabulated-curve interpolation entirely. When target_frequencies is omitted for an array target it defaults to the one-sided grid matching its length at the configured sampling frequency.

Parameters:

Name Type Description Default
psd_file str | Path | None

Path to a two-column PSD table.

None
target_psd ndarray | None

One-sided target PSD array, mutually exclusive with psd_file.

None
target_frequencies ndarray | None

Frequencies of target_psd in hertz.

None
filter_length int

Truncation control L_f in samples.

DEFAULT_FILTER_LENGTH
minimum_phase bool

Whether to apply the cepstral minimum-phase factorisation.

True
detectors list[str] | None

Detector names to generate.

None
sampling_frequency float

Sampling frequency in hertz.

4096.0
duration float

Nominal generation duration in seconds.

4.0
seed int | None

Base seed for the per-detector generators.

None
low_frequency_cutoff float

Lower edge of the target band in hertz.

2.0
high_frequency_cutoff float | None

Upper edge of the target band in hertz.

None
block_size int

Overlap-save block length in samples.

DEFAULT_BLOCK_SIZE

continuation_state()

Return the bounded state needed to keep a running stream going.

The state is the filter memory of every detector plus the pending output and block bookkeeping. It contains no generated strain and its size is independent of the generated span.

export_state()

Return a picklable snapshot sufficient to resume the stream.

from_component(component, config) classmethod

Construct an overlap-save FIR simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector coloured noise with continuity across calls.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield coloured-noise chunks lazily while preserving filter memory.

import_state(state)

Restore a snapshot produced by :meth:export_state.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same configuration.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator's settings.

reset()

Clear the filter memory, pending samples and random-number generators.

resume_metadata()

Return the metadata a stopped stream must persist to resume.

This is the per-detector bit-generator state plus the settings needed to rebuild the filter, and is reported separately from the continuation state.

ParallelAdapter

Wrap an independent-detector simulator factory with parallel execution.

metadata property

Return metadata describing the wrapped simulator and executor.

segment_gps_start property writable

Return the GPS epoch assigned for the next segment, or None if never set.

__init__(base_factory, *, max_workers=None, backend='auto')

Store the simulator factory and executor preferences.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector strain arrays in parallel.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield parallel-generated chunks lazily.

reset()

Clear cached worker simulators, resetting continuity across workers.

PowerLawSNRDistribution dataclass

Bases: SNRDistribution

Power-law target SNR above a threshold.

With maximum unset the survival function above minimum goes as S(s) = (s / minimum) ** -alpha, which is the convention a Hill or maximum-likelihood tail index is quoted in: a class measured to have a tail index of 1.34 above SNR 10 is configured as {distribution = "power_law", minimum = 10.0, alpha = 1.34} with no conversion. Note that alpha is the survival exponent, one less than the density's. Setting maximum renormalizes that survival onto [minimum, maximum] as (S(s) - S(maximum)) / (1 - S(maximum)), which reaches zero at the cap rather than continuing past it.

minimum is a threshold, not a fit to the whole population: the measured index describes the tail above it and says nothing about the bulk below, so a model configured this way reproduces the tail and replaces the bulk with the same power law continued down to the threshold.

maximum truncates the draw. It is optional and unset by default, but for a heavy tail it is worth setting: with alpha below 1 the untruncated distribution has no finite mean, and a long enough run will eventually draw an SNR no detector could produce. The largest SNR observed for the class is the natural choice.

Below alpha of about 0.052 it stops being optional and is required: the untruncated draw then runs off the top of the float range (see :func:_untruncated_headroom_bits), which a configuration cannot express and a run cannot use, so such a configuration is refused at construction rather than left to fail partway through a run. Every index in the measured range -- 0.40 for the heaviest Gravity Spy class, 3.11 for the lightest -- is far above that, and is unaffected.

Attributes:

Name Type Description
minimum float

Threshold SNR; every draw is at or above it.

alpha float

Tail index of the survival function. Must be greater than zero.

maximum float | None

Optional upper truncation, strictly above minimum.

distribution str

Discriminator naming this shape in a configuration.

__post_init__()

Validate the configured power-law parameters.

sample(rng)

Draw one target SNR by inverting the survival function.

rng.random() returns a uniform on [0, 1), so the inverted survival minimum * (1 - u) ** (-1 / alpha) covers [minimum, inf) and never divides by zero. Truncation rescales the same inversion onto [minimum, maximum].

The untruncated branch is guarded rather than trusted. __post_init__ already refuses a configuration whose draws can leave the float range, so nothing built through it reaches the guard; what the guard covers is a value that got past that check anyway -- an instance mutated after construction, or a configuration sitting on the boundary where the bound is computed in logarithms and can be a rounding hair out. An unusable target has to stop here either way, because past this point it is a scale factor on a waveform and a number in a truth catalogue, where infinity is indistinguishable from a very loud glitch.

serialize()

Return the mapping that reconstructs this distribution.

survival(threshold)

Return the fraction of draws at or above threshold.

The selection factor of a recording or detection cut applied to this class: how much of the configured population a search at that threshold would ever see. Read off the distribution in closed form rather than estimated from a realization, so an expected-count decomposition is exact rather than Monte-Carlo.

Parameters:

Name Type Description Default
threshold float

The SNR cut.

required

Returns:

Type Description
float

One at or below minimum, zero above maximum where the tail is truncated,

float

and the survival function in between.

SNRDistribution dataclass

Base class for a sampled target-SNR distribution.

sample(rng)

Draw one target SNR.

serialize()

Return the mapping that reconstructs this distribution.

survival(threshold)

Return the fraction of draws at or above threshold.

ScatteredLightGlitch dataclass

Bases: GlitchModel

Arch-shaped scattered-light transient with a Gaussian envelope.

__post_init__()

Validate scattered-light parameters.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate a chirping scattered-light glitch.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

SchumannNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate correlated strain noise from an isotropic Schumann-resonance model.

metadata property

Return simulator metadata.

previous_strain property

Expose continuity buffers for protocol-compatible state inspection.

__init__(*, positions, coupling_files, detectors=None, schumann_params=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, window_duration=DEFAULT_WINDOW_DURATION)

Initialize the Schumann simulator.

from_component(component, config) classmethod

Construct a Schumann-noise simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate correlated Schumann strain noise.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield Schumann-noise chunks lazily while preserving simulator state.

reset()

Clear continuity and RNG state.

spectrum(frequencies)

Return the isotropic Schumann magnetic PSD on a frequency grid.

theoretical_coherence(frequency, detector_a, detector_b)

Return the isotropic Schumann coherence approximation.

SchumannParams dataclass

Physical parameters for the isotropic Schumann-resonance model.

__post_init__()

Validate the resonance parameter vectors.

SimulationResult dataclass

Result of a noise simulation run.

Attributes:

Name Type Description
output_paths dict[str, Path]

Paths to generated output files, keyed by detector name.

config NoiseConfig

The configuration used for the simulation.

SpectralCovariance dataclass

Spectral covariance arrays on a common frequency grid.

SpectralLine dataclass

Configuration for a narrow-band spectral line.

__post_init__()

Validate scalar line parameters.

SpectralLineSimulator

Bases: ConfigurableNoiseSimulator

Generate additive spectral lines directly in the time domain.

metadata property

Return simulator metadata.

__init__(*, lines, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None)

Initialize the line generator.

from_component(component, config) classmethod

Construct a spectral-line simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate line-only strain for all requested detectors.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield spectral-line chunks lazily.

reset()

Clear continuity state and phase initialization.

WhiteNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate Gaussian white noise as a composable component.

metadata property

Return metadata describing the white-noise component.

__init__(*, duration=4.0, sampling_frequency=4096.0, detectors=None, seed=None)

Initialize the white-noise component state.

from_component(component, config) classmethod

Construct one white-noise component from generic config.

generate(duration, sampling_frequency, detectors, seed=None)

Return Gaussian white-noise strain arrays.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield white-noise strain chunks lazily.

__getattr__(name)

Lazily resolve optional top-level exports.

apply_segment_gps_start(simulator, gps_start)

Point every glitch injector inside a simulator chain at one segment epoch.

A writer that assigns a per-segment epoch -- FrameWriter, whose frame names carry it -- is the authority on where a segment sits in GPS time, and the glitch truth catalogue has to be stamped with that same instant or the catalogue and the frame name describe different data. Threading the writer's epoch through is what keeps it one number rather than two that agree only while the segments happen to be contiguous.

Walks the wrapper chain (base) and a composite's components, because the injector is rarely the outermost simulator: a glitch component sits inside CompositeNoiseSimulator, which sits inside whatever adapter the writer holds. A simulator with no injector anywhere beneath it is left alone.

A wrapper that builds its simulators rather than holding one -- ParallelAdapter, which constructs and caches one per detector from a factory -- cannot be reached by walking attributes, and an epoch set on such a wrapper alone reached none of its workers. Those declare a writable segment_gps_start and own the fan-out, so this hands the epoch over and lets them place both the workers they already hold and any they build later in the same segment.

assemble_hermitian_spectral_matrices(detectors, psd, csd=None)

Assemble one Hermitian spectral covariance matrix per frequency.

available_glitch_populations()

Return the registered population names, sorted.

available_joint_backend_names()

Return the registered joint backend names, sorted.

build_spectral_covariance_from_files(*, detectors, psd_files, csd_files, frequencies, taper, delta_frequency, regularization_epsilon=1e-12)

Load PSD/CSD files and build spectral covariance products.

cholesky_factors_from_spectral_matrices(spectral_matrices, *, delta_frequency, regularization_epsilon=1e-12)

Build coefficient-covariance Cholesky factors from one-sided spectra.

compare_psd(data, target_psd_file, sampling_frequency, rtol=0.1, fmin=10.0, fmax=None)

Compare an estimated PSD against a target PSD file over a frequency band.

This helper compares the median PSD level in-band rather than every bin individually, which makes it robust to the residual variance of a single Welch estimate from stochastic noise.

discover_joint_backends() cached

Return the entry points registered under :data:JOINT_BACKEND_ENTRY_POINT_GROUP.

This performs discovery only -- it does not import any backend's module. Call :meth:importlib.metadata.EntryPoint.load on a returned entry point (or use :func:load_joint_backend) to import and resolve the backend class it names.

estimate_psd(data, sampling_frequency, segment_duration=4.0, overlap=0.5)

Estimate a one-sided PSD in physical density units.

The returned frequencies are in Hz. For strain input data, the PSD is in strain^2/Hz.

et_o3_anchored_population(*, psd_file=DEFAULT_PSD, participation_probability=COHERENT_PARTICIPATION_PROBABILITY, amplitude_ratio_std=COHERENT_AMPLITUDE_RATIO_STD, coherent_rate_hz=COHERENT_RATE_HZ, short_rate_hz=SHORT_RATE_HZ)

Build the et-o3-anchored-v1 population.

Called with no arguments this returns the registered population, and its digest is the one a campaign pins. The arguments exist so the assumptions that could not be anchored can be varied deliberately -- a sweep over the participation probability is the obvious one -- and any such variation produces a different digest, which is the point: a run cannot claim to have used the registered population while having weakened it.

Note that short_rate_hz and coherent_rate_hz are the rates of two different kinds of process. The short class runs one Poisson process per interferometer, so its rate is what each of them sees. The longer class runs one process for the whole network, so its rate is the rate of shared events; each interferometer sees that times the participation probability.

Parameters:

Name Type Description Default
psd_file str | Path

Noise curve the target SNRs are calibrated against. A bundled preset name keeps the digest portable; an absolute path does not.

DEFAULT_PSD
participation_probability float

Probability that one interferometer receives a given event of the coherent class. Unanchored; see :data:UNANCHORED.

COHERENT_PARTICIPATION_PROBABILITY
amplitude_ratio_std float

Spread of the per-interferometer amplitude ratio of the coherent class. Unanchored.

COHERENT_AMPLITUDE_RATIO_STD
coherent_rate_hz float

Network event rate of the coherent class, in Hz.

COHERENT_RATE_HZ
short_rate_hz float

Per-interferometer event rate of the short class, in Hz.

SHORT_RATE_HZ

Returns:

Type Description
GlitchPopulation

The population, with both classes scoped to every interferometer in the run.

get_glitch_population(name)

Return a registered population by name.

Each call builds a fresh population, so a caller cannot mutate the registry's copy out from under the next caller. Two calls compare equal and digest identically.

Parameters:

Name Type Description Default
name str

The registered name, e.g. "et-o3-anchored-v1".

required

Returns:

Type Description
GlitchPopulation

The population.

Raises:

Type Description
KeyError

If no population is registered under that name. The message lists the names that are, because a typo is the likeliest cause and a bare KeyError does not say what to type instead.

interpolate_complex_spectral_series(frequencies, values, target_frequencies, *, taper=None)

Interpolate a complex spectral series by real and imaginary parts.

interpolate_real_spectral_series(frequencies, values, target_frequencies, *, taper=None)

Interpolate and clip a real non-negative spectral series.

load_and_interpolate_csd(file_path, target_frequencies, *, taper=None)

Load and interpolate a complex one-sided CSD file.

load_and_interpolate_psd(file_path, target_frequencies, *, taper=None)

Load and interpolate a one-sided PSD file.

load_config(path)

Load and validate noise configuration from a file.

Supports TOML, YAML, and JSON formats.

Parameters:

Name Type Description Default
path Path | str

Path to the configuration file.

required

Returns:

Type Description
NoiseConfig

Validated NoiseConfig instance.

Raises:

Type Description
FileNotFoundError

If the config file does not exist.

ValueError

If the file format is unsupported or validation fails.

load_joint_backend(name)

Import and return the backend class registered under name.

Parameters:

Name Type Description Default
name str

The entry-point name a package registered its backend under.

required

Returns:

Type Description
Any

The backend class (or factory callable) the entry point names.

Raises:

Type Description
KeyError

If no backend is registered under name.

normalize_csd_mapping(csd_values, *, detectors)

Validate and normalize detector-pair CSD mappings.

normalize_detector_pair(pair)

Normalize a detector-pair key to sorted order.

open_stream(simulator, *, chunk_duration, sampling_frequency, detectors, seed=None)

Open a stateful chunk stream on a protocol-compatible simulator.

regularized_cholesky(spectral_matrix, *, regularization_epsilon=1e-12)

Return a numerically stable Cholesky factor for a Hermitian matrix.

run_diagnostics(data, sampling_frequency, whiten=True)

Run the default Gaussianity and stationarity diagnostics.

Data is whitened against its own Welch PSD estimate before testing (see _whiten_samples), which makes both checks applicable to coloured noise. Pass whiten=False to test the raw samples instead — only meaningful for broadband data.

sample_complex_frequency_coefficients(rng, cholesky_factors, *, real_only_indices=None)

Sample complex Gaussian frequency coefficients from Cholesky factors.

simulate_spectral_covariance_chunk(rng, cholesky_factors, *, detectors, frequency_grid_size, frequency_mask, delta_frequency, window_size)

Sample one real multi-detector chunk from spectral covariance factors.

take(stream, total_duration, chunk_duration, sampling_frequency)

Collect stream chunks up to total_duration seconds.

time_series_from_frequency_coefficients(coefficients, *, detectors, frequency_grid_size, frequency_mask, delta_frequency, window_size)

Convert masked detector frequency coefficients to real time series.

Configuration

Pydantic schema models and load_config for TOML/YAML/JSON.

Configuration schema and loading helpers for gwmock_noise.

NoiseComponentConfig

Bases: BaseModel

One configurable noise component in a composed simulation.

normalize_component_definition(value) classmethod

Accept string, flat mapping, or explicit {simulator, options} input.

validate_component()

Validate the normalized component entry.

NoiseConfig

Bases: BaseModel

Generic configuration for composed detector-noise simulations.

OutputConfig

Bases: BaseModel

Configuration for simulation output.

load_config(path)

Load and validate noise configuration from a file.

Supports TOML, YAML, and JSON formats.

Parameters:

Name Type Description Default
path Path | str

Path to the configuration file.

required

Returns:

Type Description
NoiseConfig

Validated NoiseConfig instance.

Raises:

Type Description
FileNotFoundError

If the config file does not exist.

ValueError

If the file format is unsupported or validation fails.

Gaussian noise models and helpers

Spectral-line definitions plus PSD reference helpers used by Gaussian-noise configuration and simulators.

Gaussian-noise domain models and helpers.

SpectralLine dataclass

Configuration for a narrow-band spectral line.

__post_init__()

Validate scalar line parameters.

is_remote_psd_reference(value)

Return whether a PSD reference is an HTTP(S) URL.

normalize_spectral_lines(value)

Normalize heterogeneous spectral-line inputs to dataclass instances.

resolve_bundled_psd_preset(value)

Resolve a bare preset name to a bundled PSD asset.

Spectral covariance utilities

Signal-agnostic helpers for loading PSD/CSD files, interpolating spectra onto an FFT grid, assembling Hermitian covariance matrices, regularizing Cholesky factors, and sampling real multi-detector time series.

Signal-agnostic spectral covariance utilities.

The utilities in this module use one-sided PSD/CSD inputs in units of strain^2/Hz and sample masked real-FFT coefficients with covariance S(f) / (2 df). The inverse transform multiplies by df * n so the generated real time series is consistent with the one-sided periodogram convention used by the simulators.

SpectralCovariance dataclass

Spectral covariance arrays on a common frequency grid.

assemble_hermitian_spectral_matrices(detectors, psd, csd=None)

Assemble one Hermitian spectral covariance matrix per frequency.

build_spectral_covariance_from_files(*, detectors, psd_files, csd_files, frequencies, taper, delta_frequency, regularization_epsilon=1e-12)

Load PSD/CSD files and build spectral covariance products.

cholesky_factors_from_spectral_matrices(spectral_matrices, *, delta_frequency, regularization_epsilon=1e-12)

Build coefficient-covariance Cholesky factors from one-sided spectra.

interpolate_complex_spectral_series(frequencies, values, target_frequencies, *, taper=None)

Interpolate a complex spectral series by real and imaginary parts.

interpolate_real_spectral_series(frequencies, values, target_frequencies, *, taper=None)

Interpolate and clip a real non-negative spectral series.

load_and_interpolate_csd(file_path, target_frequencies, *, taper=None)

Load and interpolate a complex one-sided CSD file.

load_and_interpolate_psd(file_path, target_frequencies, *, taper=None)

Load and interpolate a one-sided PSD file.

normalize_csd_mapping(csd_values, *, detectors)

Validate and normalize detector-pair CSD mappings.

normalize_detector_pair(pair)

Normalize a detector-pair key to sorted order.

regularized_cholesky(spectral_matrix, *, regularization_epsilon=1e-12)

Return a numerically stable Cholesky factor for a Hermitian matrix.

sample_complex_frequency_coefficients(rng, cholesky_factors, *, real_only_indices=None)

Sample complex Gaussian frequency coefficients from Cholesky factors.

simulate_spectral_covariance_chunk(rng, cholesky_factors, *, detectors, frequency_grid_size, frequency_mask, delta_frequency, window_size)

Sample one real multi-detector chunk from spectral covariance factors.

time_series_from_frequency_coefficients(coefficients, *, detectors, frequency_grid_size, frequency_mask, delta_frequency, window_size)

Convert masked detector frequency coefficients to real time series.

Glitch models

Runtime glitch model definitions, including phenomenological models, the optional gengli-backed blip implementation, and the optional DeepExtractor model that injects real O3 glitch reconstructions.

Public glitch-model implementations.

BlipGlitch dataclass

Bases: GlitchModel

Gaussian-windowed broadband burst.

__post_init__()

Validate blip-specific parameters.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate a Gaussian-windowed white-noise burst.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

DeepExtractorGlitch dataclass

Bases: GlitchModel

Real O3 glitch reconstructions colored against a target PSD.

Waveforms come from the DeepExtractor glitch-reconstruction dataset: whitened, amplitude-normalized 2-second time series sampled at 4096 Hz, covering seven Gravity Spy classes. Each drawn waveform is resampled to the simulation rate, colored against the configured PSD, and rescaled so its optimal SNR sqrt(4 df sum(|h(f)|^2 / S(f))) against that PSD matches the configured target. The dataset (~2.3 GB) is downloaded lazily on first use and cached by huggingface_hub; later runs reuse the cached files. On each run the Hub is contacted first to validate the cached files' ETags and download anything missing. If the Hub is unreachable, a warning notes that the ETag check was skipped and the cached files are used instead; if they are not cached, the first generate_waveform raises huggingface_hub's LocalEntryNotFoundError (a FileNotFoundError subclass). Set local_files_only=True to skip the network unconditionally and read straight from the cache.

revision pins the download to a specific dataset version — a git branch, tag, or commit SHA forwarded to hf_hub_download. Leave it None to track the repository default. Either way, the first download resolves to a concrete commit SHA (read back from the Hub cache path), which is what serialize records. Replaying a run through its metadata therefore fetches the exact commit that produced it, keeping glitch generation bit-reproducible for a fixed (version, config, seed) even as the upstream dataset moves.

rate accepts either a single number — the total Poisson rate shared by all configured classes, drawn uniformly — or a mapping from class name to per-class Poisson rate, in which case the total rate is their sum and each event's class is drawn proportionally to its rate.

snr accepts a number, a mapping from class name to number, a distribution, or a mapping from class name to distribution — freely mixed, so one class can be sampled while another stays fixed. A number calibrates every event of that class to exactly that loudness; a distribution draws a target per event from the model's own random stream, which is what a measured glitch population needs, since each Gravity Spy class is heavy-tailed and the tail indices differ by close to an order of magnitude between classes. See :mod:gwmock_noise.glitches.snr for the supported shapes.

What the target means is the caller's to decide. A tail index measured from a trigger generator's SNR — Omicron's, say — is defined against that pipeline's PSD over the band it searched, while this model calibrates against psd_file from low_frequency_cutoff upward. Configuring one from the other equates two SNRs defined against different noise curves over different bands. The sampler reproduces whatever distribution it is given and cannot tell whether that identification is the one you meant.

Resampling uses linear interpolation without an anti-aliasing filter, so sampling frequencies below 4096 Hz alias high-frequency content; the SNR calibration itself is unaffected because it is computed after resampling.

__post_init__()

Validate the configuration and preload the PSD table.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate one colored, SNR-calibrated DeepExtractor glitch.

resolve()

Pin the dataset to a concrete commit SHA, fetching only the samples file.

Reads the commit SHA huggingface_hub stamps into the cache path so the resolved revision is available for reproducibility metadata even when no waveform has been drawn yet — e.g. a batch that fires zero glitch events, where generate_waveform (and therefore _get_dataset) never runs. The samples file is downloaded once and reused by the later full load, so this adds no extra download on the generation path. Returns the resolved SHA, or None when the cache layout does not expose one (e.g. a test stub serving files from a flat directory).

serialize()

Return metadata-friendly model parameters.

EmpiricalSNRDistribution dataclass

Bases: SNRDistribution

Target SNR drawn with replacement from a table of observed SNRs.

For a caller holding the measurements rather than a fit to them: the drawn population is exactly the supplied one, tail included, with no assumption about its shape. Supply the values inline as samples or point file at a table on disk -- exactly one of the two.

Attributes:

Name Type Description
samples Sequence[float] | None

Observed SNRs given inline, or None when file is used.

file str | Path | None

Path to a table of observed SNRs, or None when samples is used. Read by :func:load_snr_samples.

distribution str

Discriminator naming this shape in a configuration.

__post_init__()

Validate the configuration and load the SNR table.

sample(rng)

Draw one target SNR uniformly from the table, with replacement.

serialize()

Return the mapping that reconstructs this distribution.

Echoes whichever of the two inputs was configured, so a file-backed table replays as a reference to the file rather than as a copy of its contents inlined into the metadata.

survival(threshold)

Return the fraction of the supplied table at or above threshold.

GengliBlipGlitch dataclass

Bases: GlitchModel

File-backed gengli blip generator colored against a target PSD.

__post_init__()

Validate configured files and preload the population/PSD tables.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

from_population_file(population_file, **kwargs) classmethod

Construct a gengli glitch model from a population file path.

generate_waveform(sampling_frequency, rng=None)

Generate one colored gengli blip waveform.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

GlitchDraw

Bases: NamedTuple

One drawn glitch waveform and the parameters that produced it.

What :meth:GlitchModel.draw returns so an injector can record what it injected, not merely that it injected something. Every field but the waveform is optional because not every model has the quantity: a parametric blip with no PSD has no SNR to report, and only the dataset-backed models draw a Gravity Spy class.

Attributes:

Name Type Description
waveform ndarray

The glitch strain, sampled at the requested rate. Its first sample lands at the injection time -- the waveform starts there, it does not peak there.

amplitude float | None

The amplitude multiplier drawn from the model's amplitude distribution, or None when the model does not draw one.

glitch_class str | None

The drawn morphological class (e.g. a Gravity Spy class name), or None for a model with a single morphology.

target_snr float | None

The optimal SNR the draw was calibrated to before the amplitude multiplier, or None when the model does not calibrate SNR. This is the configured target, not what the strain holds.

realized_snr float | None

The optimal SNR of the whole of waveform against the PSD it was colored with, or None for an uncolored model. Only the end of the generated data can leave the strain holding less than this; a waveform crossing a segment boundary has its remainder injected into the next segment.

How it relates to target_snr is model-specific, so do not assume one from the other. A model that calibrates against its own coloring PSD -- BlipGlitch, ScatteredLightGlitch, DeepExtractorGlitch -- realizes amplitude * target_snr. GengliBlipGlitch does not: its target is sampled from a population and imposed by gengli on the whitened waveform, while the realized figure is measured on the colored, amplitude-scaled result, so the two are independent numbers rather than one restated.

GlitchModel dataclass

Base dataclass for transient glitch generators.

detectors scopes the model to a subset of the run's interferometers. It accepts one name or a list of them, and defaults to None, which is every interferometer -- what a model without a selector has always meant. Scoping is what lets one configuration carry a 10 km model for one set of interferometers and a 15 km model for another, and per-interferometer rates for a network whose instruments are not equally glitchy (in O3, Fast_Scattering fired about 29 times more often in L1 than in H1).

A selector is only half of it. :func:validate_glitch_detector_coverage refuses a run whose models do not between them account for every interferometer, or whose models color one interferometer against two different PSDs -- see its docstring for why proceeding was the defect.

__post_init__()

Validate common glitch parameters.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate a single glitch waveform.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

LogNormalAmplitudeDistribution dataclass

Log-normal amplitude sampler parameterized by linear mean and std.

__post_init__()

Validate the configured distribution parameters.

sample(rng)

Draw one amplitude sample.

NetworkCoherence dataclass

Share one glitch model's events across the interferometers it applies to.

Without it, every glitch model runs an independent Poisson process and an independent random stream per interferometer, so two interferometers never carry the same transient. That is the right default, and it is what the one measurement of the question says for widely separated sites: over O1 and O2 the LIGO blip population produced no coincidences inside the ±15 ms window an astrophysical signal can occupy, and the coincidences found in wider windows matched the accidental expectation.

It is not what a network of co-located interferometers is expected to do. Three interferometers sharing one site, one vacuum system and one seismic environment see a common environmental transient in all three at once, and a coincidence veto -- or a null stream -- assumes exactly the incoherence that such an event breaks. A model carrying this runs one Poisson process for the whole network, draws one waveform per event, and offers that waveform to each interferometer it applies to.

Participation is decided per interferometer, independently, and is never conditioned on how many others took the event. Conditioning -- "keep only events that land in at least two interferometers" -- is the obvious way to write a coincident population, and it is the wrong one here: it makes what one interferometer's strain contains depend on which other interferometers the run happens to include, so the same channel stops being reproducible between a three-interferometer run and a two-interferometer one. Paired-geometry comparisons need that reproducibility, so an event that lands nowhere simply lands nowhere, and the multiplicity comes out Binomial rather than being imposed.

Attributes:

Name Type Description
participation_probability float

Probability that any one interferometer the model applies to receives a given network event. 1.0 -- the default -- puts every event in every one of them, which is the strongest coherence the model can express.

amplitude_ratio_std float

Standard deviation of a per-interferometer lognormal amplitude ratio with linear mean 1.0, applied to the shared waveform. 0.0 -- the default -- injects the identical strain into every participating interferometer.

__post_init__()

Validate the configured coherence parameters.

Raises:

Type Description
ValueError

If the participation probability is outside [0, 1] or the amplitude-ratio spread is negative.

amplitude_ratio(rng)

Draw one interferometer's amplitude ratio for a shared waveform.

Parameters:

Name Type Description Default
rng Generator

The generator for this (event, interferometer) pair.

required

Returns:

Type Description
float

The multiplier to apply to the shared waveform.

serialize()

Return the mapping that reconstructs this coherence specification.

PowerLawSNRDistribution dataclass

Bases: SNRDistribution

Power-law target SNR above a threshold.

With maximum unset the survival function above minimum goes as S(s) = (s / minimum) ** -alpha, which is the convention a Hill or maximum-likelihood tail index is quoted in: a class measured to have a tail index of 1.34 above SNR 10 is configured as {distribution = "power_law", minimum = 10.0, alpha = 1.34} with no conversion. Note that alpha is the survival exponent, one less than the density's. Setting maximum renormalizes that survival onto [minimum, maximum] as (S(s) - S(maximum)) / (1 - S(maximum)), which reaches zero at the cap rather than continuing past it.

minimum is a threshold, not a fit to the whole population: the measured index describes the tail above it and says nothing about the bulk below, so a model configured this way reproduces the tail and replaces the bulk with the same power law continued down to the threshold.

maximum truncates the draw. It is optional and unset by default, but for a heavy tail it is worth setting: with alpha below 1 the untruncated distribution has no finite mean, and a long enough run will eventually draw an SNR no detector could produce. The largest SNR observed for the class is the natural choice.

Below alpha of about 0.052 it stops being optional and is required: the untruncated draw then runs off the top of the float range (see :func:_untruncated_headroom_bits), which a configuration cannot express and a run cannot use, so such a configuration is refused at construction rather than left to fail partway through a run. Every index in the measured range -- 0.40 for the heaviest Gravity Spy class, 3.11 for the lightest -- is far above that, and is unaffected.

Attributes:

Name Type Description
minimum float

Threshold SNR; every draw is at or above it.

alpha float

Tail index of the survival function. Must be greater than zero.

maximum float | None

Optional upper truncation, strictly above minimum.

distribution str

Discriminator naming this shape in a configuration.

__post_init__()

Validate the configured power-law parameters.

sample(rng)

Draw one target SNR by inverting the survival function.

rng.random() returns a uniform on [0, 1), so the inverted survival minimum * (1 - u) ** (-1 / alpha) covers [minimum, inf) and never divides by zero. Truncation rescales the same inversion onto [minimum, maximum].

The untruncated branch is guarded rather than trusted. __post_init__ already refuses a configuration whose draws can leave the float range, so nothing built through it reaches the guard; what the guard covers is a value that got past that check anyway -- an instance mutated after construction, or a configuration sitting on the boundary where the bound is computed in logarithms and can be a rounding hair out. An unusable target has to stop here either way, because past this point it is a scale factor on a waveform and a number in a truth catalogue, where infinity is indistinguishable from a very loud glitch.

serialize()

Return the mapping that reconstructs this distribution.

survival(threshold)

Return the fraction of draws at or above threshold.

The selection factor of a recording or detection cut applied to this class: how much of the configured population a search at that threshold would ever see. Read off the distribution in closed form rather than estimated from a realization, so an expected-count decomposition is exact rather than Monte-Carlo.

Parameters:

Name Type Description Default
threshold float

The SNR cut.

required

Returns:

Type Description
float

One at or below minimum, zero above maximum where the tail is truncated,

float

and the survival function in between.

SNRDistribution dataclass

Base class for a sampled target-SNR distribution.

sample(rng)

Draw one target SNR.

serialize()

Return the mapping that reconstructs this distribution.

survival(threshold)

Return the fraction of draws at or above threshold.

ScatteredLightGlitch dataclass

Bases: GlitchModel

Arch-shaped scattered-light transient with a Gaussian envelope.

__post_init__()

Validate scattered-light parameters.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate a chirping scattered-light glitch.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

load_snr_samples(path)

Load a one-dimensional table of observed SNRs from disk.

.h5/.hdf5 files are read from their snr dataset -- the schema gwmock-noise build-blip-glitch-table writes -- and anything else is read as a whitespace-separated text column.

Parameters:

Name Type Description Default
path str | Path

Path to the SNR table.

required

Returns:

Type Description
ndarray

The validated SNR samples.

Raises:

Type Description
FileNotFoundError

If path does not exist.

ValueError

If the table is empty, not one-dimensional, or holds a non-finite or non-positive SNR.

normalize_glitch_models(value)

Normalize heterogeneous glitch-config inputs.

Entries carry an optional detectors selector -- one interferometer name or a list of them -- which scopes the model to those interferometers; an entry without one applies to every interferometer in the run, as it always has. Whether the resulting list accounts for the run is :func:validate_glitch_detector_coverage's question, because it is the run that supplies the interferometers.

Parameters:

Name Type Description Default
value Any

The configured list of models or model mappings.

required

Returns:

Type Description
list[GlitchModel]

The normalized glitch models.

Raises:

Type Description
TypeError

If the value is not a list, or an entry is not a mapping.

ValueError

If an entry names an unsupported kind or omits its amplitude distribution.

normalize_snr(value, parameter='snr')

Normalize one target-SNR specification into a number or a distribution.

Accepts what a configuration can hold for a single target: a number, a already-constructed :class:SNRDistribution, or a mapping describing one.

Parameters:

Name Type Description Default
value Any

The configured specification.

required
parameter str

Name of the configuration key, used in error messages.

'snr'

Returns:

Type Description
float | SNRDistribution

The target SNR as a float, or the distribution to draw it from.

Raises:

Type Description
TypeError

If value is neither a number nor a distribution.

ValueError

If a numeric target is not finite and positive.

snr_survival(specification, threshold)

Return the fraction of one class's events whose target SNR reaches threshold.

The selection factor in an expected-count decomposition. A model with no target SNR at all has nothing to select on, and reports 1.0 rather than guessing: its loudness is set by the amplitude distribution in units the threshold is not expressed in.

Parameters:

Name Type Description Default
specification float | SNRDistribution | None

A fixed target SNR, a distribution to draw it from, or None.

required
threshold float

The SNR cut.

required

Returns:

Type Description
float

The surviving fraction, between zero and one.

supported_glitch_kinds()

Return all supported glitch-model kinds.

validate_glitch_detector_coverage(models, detectors)

Refuse a glitch configuration that does not say what every interferometer gets.

A glitch model used to apply to every interferometer in the run because there was no way to say otherwise, so one psd_file colored all of them. A run naming both Einstein Telescope designs resolves to interferometers of two arm lengths, and the 10 km curve was applied to the 15 km instruments as well: SNRs 23-44% away from what the configuration asked for, morphology-dependent, with nothing in the output or the logs to say so. The silence was the defect, so a configuration that cannot be read one way is refused here rather than run.

Three things are refused, in the order a reader can act on them:

  1. A selector naming an interferometer the run does not have. Usually a typo or a leftover from another network, and the model then never fires -- a configured rate and PSD that inject nothing, which nothing in the output distinguishes from a model that fired and drew no events.
  2. An interferometer no model claims. Its strain is written with no glitches while its neighbours carry them. Where that is the intent, say it: a model scoped to it with rate = 0.0 states "this interferometer carries no glitches" in the configuration, where a reader can see it.
  3. Two coloring PSDs claiming one interferometer. An interferometer has one noise floor; two models coloring it against different curves disagree about what instrument it is, and whichever ran second would not make that true. Models that do no coloring impose no floor and are not part of this.

Parameters:

Name Type Description Default
models Sequence[GlitchModel]

The run's glitch models, already normalized.

required
detectors Sequence[str]

The interferometers the run will generate.

required

Raises:

Type Description
ValueError

If any of the three cases above holds. The message names the interferometers involved.

Glitch populations

Registered glitch populations: a named, serializable, digestible set of glitch classes, with an expected-count decomposition and a reusable seeded realization that paired arms of a comparison share.

Registered glitch populations and the machinery that pins them.

ExpectedCount dataclass

The expected injected count for one class in one interferometer, factorized.

The product is :attr:expected, and every factor is carried beside it. A campaign that finds the total surprising can then read off which factor is surprising -- a rate transplanted from another run, a livetime that is calendar time rather than observing time, a participation probability nobody meant to set -- instead of re-deriving the whole chain.

Attributes:

Name Type Description
glitch_class str

Name of the class within its population.

detector str

The interferometer this row is for.

rate_hz float

The class's Poisson rate. For a class whose interferometers glitch independently this is the rate each of them sees; for a network-coherent class it is the rate of the shared network process, which is not the same number unless every interferometer takes every event.

livetime_seconds float

The analysed time this count is over.

participation float

Probability that this interferometer receives a given event of the class. One for a class with no network coherence.

selection float

Fraction of the class's events that survive the SNR cut the count was asked for. One when no cut was given.

expected float

rate_hz * livetime_seconds * participation * selection.

network_events property

Expected number of network events of this class over the same livetime.

The same arithmetic with the participation factor divided out: the events the class's process produced, rather than the subset this interferometer received. For a class with no network coherence the two coincide, because every event it draws for an interferometer is that interferometer's.

GlitchClass dataclass

One named class within a registered population.

The prose fields are part of the registered definition and therefore part of the digest, deliberately. A population whose numbers are unchanged but whose meaning has been rewritten is a different population, and a digest that did not move would say otherwise.

Attributes:

Name Type Description
name str

Stable identifier of the class within its population.

morphology str

Short slug naming the waveform family, e.g. "gaussian-windowed-broadband-burst".

description str

One sentence saying what the class represents and what it is drawn from.

model GlitchModel

The glitch model that generates it, already normalized.

__post_init__()

Validate the class definition.

Raises:

Type Description
ValueError

If the name, morphology or description is blank.

deserialize(payload) classmethod

Rebuild a class from :meth:serialize output.

Parameters:

Name Type Description Default
payload dict[str, Any]

The serialized class.

required

Returns:

Type Description
GlitchClass

The reconstructed class.

Raises:

Type Description
KeyError

If a required key is missing.

serialize()

Return the mapping that reconstructs this class.

GlitchPopulation dataclass

A named, versioned, digestible set of glitch classes.

The digest is only as portable as what goes into it. A model's psd_file is serialized as written, so a population built against /scratch/me/et.txt carries that path into its digest and no other machine can reproduce it. Name a bundled preset, or a path relative to the campaign's own tree, and the digest travels.

Attributes:

Name Type Description
name str

Stable registered identifier, e.g. "et-o3-anchored-v1".

description str

One paragraph saying what the population is for and what it is anchored to.

classes tuple[GlitchClass, ...]

The classes, in registration order.

references tuple[str, ...]

Published sources the registered quantities are anchored to, one citation per entry.

unanchored tuple[str, ...]

Quantities that could not be anchored to a published measurement, one per entry, each saying what it is and why. Carried in the population rather than only in its documentation, so a consumer reading a pinned population reads the caveats with it. An empty tuple is a claim that everything is anchored, so it is a claim to be able to defend.

schema_version str

Serialization schema, part of the digest.

__post_init__()

Validate the population definition.

Raises:

Type Description
ValueError

If the name or description is blank, the population has no classes, or two classes share a name.

build_injector(base, *, gps_start=0.0)

Wrap a base simulator so it also injects this population.

Parameters:

Name Type Description Default
base NoiseSimulator

The simulator producing the noise the glitches are added to.

required
gps_start float

GPS time of the first sample, stamped onto the truth catalogue.

0.0

Returns:

Type Description
InjectGlitches

The wrapped simulator.

canonical_json()

Return the exact bytes the digest is taken over.

Sorted keys, no insignificant whitespace, ASCII only and no NaN: a mapping that two processes can be relied on to render identically. Exposed rather than hidden so a campaign that disagrees about a digest can diff what was hashed instead of guessing.

Returns:

Type Description
str

The canonical serialization.

deserialize(payload) classmethod

Rebuild a population from :meth:serialize output.

Parameters:

Name Type Description Default
payload dict[str, Any]

The serialized population.

required

Returns:

Type Description
GlitchPopulation

The reconstructed population, which serializes and digests identically.

Raises:

Type Description
KeyError

If a required key is missing.

ValueError

If the payload declares a schema version this package cannot read.

digest()

Return the population's sha256:... digest.

What a campaign manifest pins. Every registered quantity, every prose field and the schema version go into it, so a population that has been edited cannot present itself as the one a finished run used.

Returns:

Type Description
str

The digest, algorithm-prefixed.

expected_counts(*, livetime_seconds, detectors, snr_threshold=None)

Decompose the expected injected count per class and interferometer.

Parameters:

Name Type Description Default
livetime_seconds float

Analysed time the count is over. This is observing time, not calendar time: the distinction is the commonest factor-of-1.3 error in a count derived from a published rate.

required
detectors Sequence[str]

The run's interferometers. Checked against the population's own scoping, so a decomposition cannot be computed for a network the population does not account for.

required
snr_threshold float | None

Optional recording or detection cut. When given, the selection factor is the fraction of each class's target-SNR distribution at or above it, read off the distribution in closed form.

None

Returns:

Type Description
list[ExpectedCount]

One row per (class, interferometer) pair, in registration order.

Raises:

Type Description
ValueError

If the livetime is not positive, no interferometers were given, or the population does not account for every interferometer.

model_for(glitch_class)

Return one class's glitch model by name.

Parameters:

Name Type Description Default
glitch_class str

The class name.

required

Returns:

Type Description
GlitchModel

The model that generates it.

Raises:

Type Description
KeyError

If the population has no such class.

models()

Return the population's glitch models in registration order.

realize(*, detectors, duration, sampling_frequency, seed, gps_start=0.0)

Generate this population's glitch strain on its own, with no noise under it.

This is the object paired arms share. A comparison whose arms differ in their signal content, their geometry, their detection threshold or their ranking statistic must not differ in their glitches, and the way to guarantee that is for every arm to add the same array rather than to re-run a generator and trust it to agree. The returned stamp carries the population digest and the seed, so an arm can record which realization it added -- and the digest is the field to pin, for the reason :class:GlitchRealization gives.

The realization is reproducible for a fixed (package version, population, seed) and does not depend on the base noise, on the other interferometers in the run, or on whether it was generated in one call or streamed -- see the module's tests, which assert each of those separately.

Parameters:

Name Type Description Default
detectors Sequence[str]

The interferometers to realize.

required
duration float

Length of the realization in seconds.

required
sampling_frequency float

Sampling frequency in Hz. Note that a class with a high_frequency_cutoff above the Nyquist frequency is refused rather than silently narrowed, so a population registered over a wide band needs a rate that can carry it.

required
seed int

The realization's seed.

required
gps_start float

GPS time of the first sample.

0.0

Returns:

Type Description
GlitchRealization

The glitch strain, its truth catalogue and the stamp identifying it.

Raises:

Type Description
ValueError

If the population does not account for every interferometer.

serialize()

Return the mapping that reconstructs this population.

GlitchRealization

Bases: NamedTuple

One seeded glitch realization and the stamp that identifies it.

Attributes:

Name Type Description
strain dict[str, ndarray]

The glitch contribution per interferometer, with nothing else in it. This is what paired arms add to their own base noise, and adding the same array to two arms is what makes their glitch content bit-identical.

events list[dict[str, Any]]

The per-event truth catalogue, in time order. See :data:gwmock_noise.simulators.glitches.GLITCH_CATALOGUE_COLUMNS.

stamp dict[str, Any]

What identifies this realization: the population's name and digest, the package version, and the generation parameters. Enough to reproduce strain and to prove another run used the same population.

The pinned value is digest, not the whole stamp. gwmock_noise_version is derived from the version-control description and therefore changes on every commit, including commits that touch nothing this population depends on; it is a provenance record of what produced a realization, and a manifest that compares it for equality would reject a rerun of the same population from a later commit. Pin population and digest, record the rest.

available_glitch_populations()

Return the registered population names, sorted.

et_o3_anchored_population(*, psd_file=DEFAULT_PSD, participation_probability=COHERENT_PARTICIPATION_PROBABILITY, amplitude_ratio_std=COHERENT_AMPLITUDE_RATIO_STD, coherent_rate_hz=COHERENT_RATE_HZ, short_rate_hz=SHORT_RATE_HZ)

Build the et-o3-anchored-v1 population.

Called with no arguments this returns the registered population, and its digest is the one a campaign pins. The arguments exist so the assumptions that could not be anchored can be varied deliberately -- a sweep over the participation probability is the obvious one -- and any such variation produces a different digest, which is the point: a run cannot claim to have used the registered population while having weakened it.

Note that short_rate_hz and coherent_rate_hz are the rates of two different kinds of process. The short class runs one Poisson process per interferometer, so its rate is what each of them sees. The longer class runs one process for the whole network, so its rate is the rate of shared events; each interferometer sees that times the participation probability.

Parameters:

Name Type Description Default
psd_file str | Path

Noise curve the target SNRs are calibrated against. A bundled preset name keeps the digest portable; an absolute path does not.

DEFAULT_PSD
participation_probability float

Probability that one interferometer receives a given event of the coherent class. Unanchored; see :data:UNANCHORED.

COHERENT_PARTICIPATION_PROBABILITY
amplitude_ratio_std float

Spread of the per-interferometer amplitude ratio of the coherent class. Unanchored.

COHERENT_AMPLITUDE_RATIO_STD
coherent_rate_hz float

Network event rate of the coherent class, in Hz.

COHERENT_RATE_HZ
short_rate_hz float

Per-interferometer event rate of the short class, in Hz.

SHORT_RATE_HZ

Returns:

Type Description
GlitchPopulation

The population, with both classes scoped to every interferometer in the run.

get_glitch_population(name)

Return a registered population by name.

Each call builds a fresh population, so a caller cannot mutate the registry's copy out from under the next caller. Two calls compare equal and digest identically.

Parameters:

Name Type Description Default
name str

The registered name, e.g. "et-o3-anchored-v1".

required

Returns:

Type Description
GlitchPopulation

The population.

Raises:

Type Description
KeyError

If no population is registered under that name. The message lists the names that are, because a typo is the likeliest cause and a bare KeyError does not say what to type instead.

Simulators

Simulator implementations, the NoiseSimulator protocol, the optional JointStrainWitnessSimulator protocol and its public JointDummyCorrelatedSimulator backend, ConfigurableNoiseSimulator, SimulationResult, and streaming helpers such as open_stream and take.

Noise simulators for gravitational wave detectors.

ARMANoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate stationary detector noise with a fitted ARMA / state-space model.

denominator_coefficients property

Return a copy of the ARMA denominator coefficients a_1 .. a_p.

metadata property

Return metadata describing the fitted ARMA model.

model_psd_curve property

Return the fitted model PSD on the band-masked design grid.

numerator_coefficients property

Return a copy of the ARMA numerator coefficients b_0 .. b_q.

resume_metadata_nbytes property

Serialized size of the resume metadata in bytes.

state_nbytes property

Serialized size of the continuation state in bytes.

target_psd_curve property

Return the band-masked target PSD on the design grid.

__init__(*, psd_file=None, target_psd=None, target_frequencies=None, ar_order=DEFAULT_AR_ORDER, ma_order=DEFAULT_MA_ORDER, line_frequencies=None, line_widths=None, detect_lines=False, max_lines=DEFAULT_MAX_LINES, line_prominence=DEFAULT_LINE_PROMINENCE, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization=DEFAULT_REGULARIZATION)

Initialize the simulator and fit the ARMA model once.

Parameters:

Name Type Description Default
psd_file str | Path | None

Path or bundled preset name of a two-column PSD table.

None
target_psd ndarray | None

One-sided target PSD array, mutually exclusive with psd_file.

None
target_frequencies ndarray | None

Frequencies of target_psd in hertz.

None
ar_order int

Autoregressive order p of the denominator.

DEFAULT_AR_ORDER
ma_order int

Moving-average order q of the numerator.

DEFAULT_MA_ORDER
line_frequencies list[float] | None

Explicit line frequencies to place poles at.

None
line_widths list[float] | None

Widths matching line_frequencies; defaults to :data:DEFAULT_LINE_WIDTH_HZ.

None
detect_lines bool

Whether to detect prominent lines from the target.

False
max_lines int

Maximum number of detected lines.

DEFAULT_MAX_LINES
line_prominence float

Minimum peak-to-median ratio for detection.

DEFAULT_LINE_PROMINENCE
detectors list[str] | None

Detector names to generate.

None
sampling_frequency float

Sampling frequency in hertz.

4096.0
duration float

Nominal generation duration in seconds.

4.0
seed int | None

Base seed for the per-detector generators.

None
low_frequency_cutoff float

Lower edge of the target band in hertz.

2.0
high_frequency_cutoff float | None

Upper edge of the target band in hertz.

None
block_size int

Number of samples generated per filter call.

DEFAULT_BLOCK_SIZE
regularization float

Relative ridge added to the zero-lag autocovariance.

DEFAULT_REGULARIZATION

continuation_state()

Return the bounded state needed to keep a running stream going.

export_state()

Return a picklable snapshot sufficient to resume the stream.

from_component(component, config) classmethod

Construct an ARMA simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector ARMA noise with continuity across calls.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield ARMA chunks lazily while preserving filter state.

import_state(state)

Restore a snapshot produced by :meth:export_state.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same configuration.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator's settings.

reset()

Clear filter state and RNGs.

resume_metadata()

Return the metadata a stopped stream must persist to resume.

ARNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate stateful detector noise from an AR model fit to a target PSD.

autoregressive_coefficients property

Return a copy of the fitted a_1 .. a_p coefficients.

metadata property

Return metadata describing the fitted AR model.

model_psd_curve property

Return the fitted model PSD on the band-masked design grid.

resume_metadata_nbytes property

Serialized size of the resume metadata in bytes.

state_nbytes property

Serialized size of the continuation state in bytes.

target_psd_curve property

Return the band-masked target PSD on the design grid.

__init__(*, psd_file, order=DEFAULT_AR_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization=DEFAULT_REGULARIZATION)

Initialize the simulator and fit the AR model once.

Parameters:

Name Type Description Default
psd_file str | Path

Path or bundled preset name of a two-column PSD table.

required
order int

Autoregressive order p.

DEFAULT_AR_ORDER
detectors list[str] | None

Detector names to generate.

None
sampling_frequency float

Sampling frequency in hertz.

4096.0
duration float

Nominal generation duration in seconds.

4.0
seed int | None

Base seed for the per-detector generators.

None
low_frequency_cutoff float

Lower edge of the target band in hertz.

2.0
high_frequency_cutoff float | None

Upper edge of the target band in hertz.

None
block_size int

Number of samples generated per recursion call.

DEFAULT_BLOCK_SIZE
regularization float

Relative ridge added to the zero-lag autocovariance before the recursion; 0 disables it.

DEFAULT_REGULARIZATION

continuation_state()

Return the bounded state needed to keep a running stream going.

The state is the recursion memory of every detector; its size is proportional to the order and independent of the generated span.

export_state()

Return a picklable snapshot sufficient to resume the stream.

from_component(component, config) classmethod

Construct an AR-noise simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector AR noise with continuity across calls.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield AR-noise chunks lazily while preserving recursion state.

import_state(state)

Restore a snapshot produced by :meth:export_state.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same configuration.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator's settings.

reset()

Clear detector state and RNGs.

resume_metadata()

Return the metadata a stopped stream must persist to resume.

AddLines

Wrap a base simulator and add spectral lines to its output.

metadata property

Return additive-wrapper metadata.

__init__(base, lines)

Initialize the additive wrapper.

generate(duration, sampling_frequency, detectors, seed=None)

Generate base noise and add the configured spectral lines.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield additive spectral-line chunks lazily.

reset()

Reset the additive line state and any resettable base state.

BaseNoiseSimulator

Bases: ABC

Abstract base class for noise simulators.

This interface is the stable API through which the upstream gwmock package interacts with gwmock_noise. Implementations must override :meth:run.

run(config) abstractmethod

Run the noise simulation with the given configuration.

Parameters:

Name Type Description Default
config NoiseConfig

Validated noise simulation configuration.

required

Returns:

Type Description
SimulationResult

Result containing paths to generated outputs and the config used.

BlipGlitch dataclass

Bases: GlitchModel

Gaussian-windowed broadband burst.

__post_init__()

Validate blip-specific parameters.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate a Gaussian-windowed white-noise burst.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

ChannelDomain

Bases: StrEnum

The sample domain a channel array is expressed in.

ChannelKind

Bases: StrEnum

The physical role a channel plays in a joint realization.

ChannelMetadata dataclass

Typed, machine-checkable description of one channel.

Attributes:

Name Type Description
name str

Channel name, unique within one joint realization.

kind ChannelKind

Whether the channel is a gravitational-wave strain channel or an auxiliary/witness channel.

unit str

Physical unit of the channel's samples (e.g. "strain" for dimensionless strain, or a physical unit string such as "m/s^2").

domain ChannelDomain

Whether the channel's array is a time series or a frequency-domain series.

sampling_frequency float

Sampling frequency in hertz.

dtype str

The numpy dtype name (e.g. "float64") the channel array is generated in.

ColoredNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate colored detector noise from an input PSD.

metadata property

Return simulator metadata.

previous_strain property

Expose continuity buffers for protocol-compatible state inspection.

__init__(*, psd_file=None, psd_schedule=None, psd_array=None, frequencies=None, delta_frequency=None, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, window_duration=DEFAULT_WINDOW_DURATION)

Initialize the simulator.

export_state()

Return a picklable snapshot of the running crossfade stream.

The snapshot persists the bit-generator state and the chunk counter of the stitcher plus the bookkeeping needed to interpolate the PSD schedule on resume. It contains no cached strain: the previous window is regenerated from the bit-generator state.

from_component(component, config) classmethod

Construct a colored-noise simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector colored noise.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield colored-noise chunks lazily while preserving simulator state.

import_state(state)

Resume the crossfade stream from an :meth:export_state snapshot.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same settings.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator.

reset()

Clear continuity and RNG state.

CompositeNoiseSimulator

Additively combine multiple component simulators.

metadata property

Return metadata describing the composed run.

__init__(components, *, detectors, duration, sampling_frequency, seed)

Initialize the composed simulator.

generate(duration, sampling_frequency, detectors, seed=None)

Generate all component realizations and add them together.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield additive component chunks lazily.

ConfigurableNoiseSimulator

Bases: ABC

Abstract mixin for built-in simulators usable as composed components.

from_component(component, config) abstractmethod classmethod

Construct one simulator instance from a generic component config.

CorrelatedARNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate correlated detector noise with a truncated VMA representation.

metadata property

Return metadata describing the fitted VMA model.

__init__(*, psd_files, csd_files=None, order=DEFAULT_VMA_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization_epsilon=DEFAULT_REGULARIZATION_EPSILON)

Initialize the simulator and fit truncated VMA taps once.

from_component(component, config) classmethod

Construct a correlated-AR simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate correlated per-detector noise with stateful continuity.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield correlated AR chunks lazily while preserving innovation state.

reset()

Clear innovation state and RNG.

CorrelatedNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate correlated detector noise from PSD and CSD inputs.

metadata property

Return simulator metadata.

previous_strain property

Expose continuity buffers for protocol-compatible state inspection.

__init__(*, psd_files, csd_files=None, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, regularization_epsilon=1e-12, window_duration=DEFAULT_WINDOW_DURATION)

Initialize the simulator.

from_component(component, config) classmethod

Construct a correlated-noise simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate correlated per-detector noise.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield correlated-noise chunks lazily while preserving simulator state.

reset()

Clear continuity and RNG state.

DefaultNoiseSimulator

Bases: BaseNoiseSimulator

Default noise simulator implementation.

metadata property

Return metadata describing the current simulator state.

__init__(*, duration=4.0, sampling_frequency=4096.0, detectors=None, seed=None)

Initialize the simulator with protocol-compatible state.

generate(duration, sampling_frequency, detectors, seed=None)

Return Gaussian white-noise strain arrays.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield white-noise strain chunks lazily.

run(config)

Run the noise simulation with the given configuration.

FitError

Bases: ValueError

Raised when a model fit is numerically invalid or unstable.

GlitchDraw

Bases: NamedTuple

One drawn glitch waveform and the parameters that produced it.

What :meth:GlitchModel.draw returns so an injector can record what it injected, not merely that it injected something. Every field but the waveform is optional because not every model has the quantity: a parametric blip with no PSD has no SNR to report, and only the dataset-backed models draw a Gravity Spy class.

Attributes:

Name Type Description
waveform ndarray

The glitch strain, sampled at the requested rate. Its first sample lands at the injection time -- the waveform starts there, it does not peak there.

amplitude float | None

The amplitude multiplier drawn from the model's amplitude distribution, or None when the model does not draw one.

glitch_class str | None

The drawn morphological class (e.g. a Gravity Spy class name), or None for a model with a single morphology.

target_snr float | None

The optimal SNR the draw was calibrated to before the amplitude multiplier, or None when the model does not calibrate SNR. This is the configured target, not what the strain holds.

realized_snr float | None

The optimal SNR of the whole of waveform against the PSD it was colored with, or None for an uncolored model. Only the end of the generated data can leave the strain holding less than this; a waveform crossing a segment boundary has its remainder injected into the next segment.

How it relates to target_snr is model-specific, so do not assume one from the other. A model that calibrates against its own coloring PSD -- BlipGlitch, ScatteredLightGlitch, DeepExtractorGlitch -- realizes amplitude * target_snr. GengliBlipGlitch does not: its target is sampled from a population and imposed by gengli on the whitened waveform, while the realized figure is measured on the colored, amplitude-scaled result, so the two are independent numbers rather than one restated.

GlitchModel dataclass

Base dataclass for transient glitch generators.

detectors scopes the model to a subset of the run's interferometers. It accepts one name or a list of them, and defaults to None, which is every interferometer -- what a model without a selector has always meant. Scoping is what lets one configuration carry a 10 km model for one set of interferometers and a 15 km model for another, and per-interferometer rates for a network whose instruments are not equally glitchy (in O3, Fast_Scattering fired about 29 times more often in L1 than in H1).

A selector is only half of it. :func:validate_glitch_detector_coverage refuses a run whose models do not between them account for every interferometer, or whose models color one interferometer against two different PSDs -- see its docstring for why proceeding was the defect.

__post_init__()

Validate common glitch parameters.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate a single glitch waveform.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

GlitchNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate transient glitches as an additive standalone component.

from_component(component, config) classmethod

Construct a glitch-only component from one component definition.

GwoscNoiseSimulator

Real-noise simulator that fetches strain data from GWOSC.

Implements the NoiseSimulator protocol so it can be used interchangeably with synthetic simulators like ColoredNoiseSimulator and CorrelatedNoiseSimulator.

Fetches real detector strain from the Gravitational-Wave Open Science Centre and applies user-configured filters to exclude segments containing GW signals or data-quality issues.

When cache_dir is set in the config, downloaded HDF5 files are saved locally and reused — avoiding repeated downloads for the same GPS interval.

Attributes:

Name Type Description
duration float

Duration of the configured GPS interval (seconds).

sampling_frequency float

Sampling frequency in Hz.

detectors list[str]

List of detector prefixes.

seed None

Always None (real noise has no random seed).

config

The underlying GWOSC configuration.

detectors property

Return the list of detector prefixes.

duration property

Return the total duration of the configured GPS interval.

metadata property

Return metadata describing the simulator and its configuration.

Returns:

Type Description
dict[str, Any]

A dictionary with the simulator implementation name,

dict[str, Any]

GPS range, sample rate, detectors, filter configuration,

dict[str, Any]

and cache status.

sampling_frequency property

Return the sampling frequency in Hz.

seed property

Return None — real noise has no controllable random seed.

__init__(config)

Initialize the real-noise simulator.

Parameters:

Name Type Description Default
config GwoscNoiseConfig

Configuration specifying GPS range, detectors, sample rate, filtering options, and optional cache directory.

required

check_availability(detectors=None)

Probe GWOSC for per-detector strain-data availability.

Convenience passthrough to :meth:GwoscNoiseFetcher.check_availability. Use this as a pre-flight check: a detector may have a fully clean interval (no vetoes) yet have no published strain data, in which case :meth:generate would fail with a clear error.

Parameters:

Name Type Description Default
detectors list[str] | None

Detectors to probe. Defaults to all configured detectors when None.

None

Returns:

Type Description
dict[str, bool]

A dictionary mapping each requested detector to True if

dict[str, bool]

GWOSC has data covering the interval, False otherwise.

generate(duration, sampling_frequency, detectors, seed=None)

Fetch clean noise and return the longest contiguous segment per detector.

Real GWOSC noise is fragmented by the excision of GW events and data-quality vetoes, so the clean data generally consists of several non-contiguous spans. To honour the NoiseSimulator contract — a single, gap-free array per detector — this returns the longest contiguous clean segment for each detector.

Concatenating the fragments instead (the historical behaviour) splices non-adjacent samples together, creating discontinuities that show up as sharp transients after whitening. Use :meth:generate_segments to obtain every clean segment separately when the full clean data is required.

Parameters:

Name Type Description Default
duration float

Requested duration (ignored; the longest available contiguous clean segment determines the output length).

required
sampling_frequency float

Requested sampling frequency (must match the configured sample_rate).

required
detectors list[str]

Requested detector list (must be a subset of the configured detectors).

required
seed int | None

Ignored for real noise.

None

Returns:

Type Description
dict[str, ndarray]

A dictionary mapping each detector to a 1-D numpy array

dict[str, ndarray]

holding its longest contiguous clean segment.

Raises:

Type Description
ValueError

If sampling_frequency does not match the configured value, if detectors are not a subset, or if a detector has no clean data.

generate_segments(duration, sampling_frequency, detectors, seed=None)

Fetch clean noise and return per-detector lists of contiguous segments.

Unlike :meth:generate, this preserves the gap structure of the real data: each clean segment — the spans left after excluding GW events and data-quality vetoes — is returned as its own array, in GPS-time order. Adjacent segments are not contiguous in time, so callers must treat each segment independently; concatenating them would splice non-adjacent samples together and introduce discontinuities at the joins (which manifest as sharp transients after whitening).

Parameters:

Name Type Description Default
duration float

Requested duration (ignored; the GPS interval from the config determines the available data).

required
sampling_frequency float

Requested sampling frequency (must match the configured sample_rate).

required
detectors list[str]

Requested detector list (must be a subset of the configured detectors).

required
seed int | None

Ignored for real noise.

None

Returns:

Type Description
dict[str, list[ndarray]]

A dictionary mapping each detector to a list of 1-D numpy

dict[str, list[ndarray]]

arrays, one per clean segment, ordered by GPS time.

Raises:

Type Description
ValueError

If sampling_frequency does not match the configured value, if detectors are not a subset, or if a detector has no clean data.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield clean-noise chunks lazily.

Fetches the longest contiguous clean segment once (see :meth:generate) and yields it in chunks of chunk_duration seconds. Chunking the longest contiguous span keeps every chunk free of the discontinuities that splicing fragments would introduce.

Parameters:

Name Type Description Default
chunk_duration float

Duration of each yielded chunk in seconds.

required
sampling_frequency float

Requested sampling frequency (must match the configured sample_rate).

required
detectors list[str]

Requested detector list (must be a subset of the configured detectors).

required
seed int | None

Ignored for real noise.

None

Yields:

Type Description
dict[str, ndarray]

Per-detector strain arrays for each chunk.

InjectGlitches

Wrap a base simulator and inject transient glitches additively.

A glitch model with no network coherence runs an independent Poisson process per detector, so every detector receives its own event times and waveform realizations, and GlitchModel.rate is the event rate seen by each individual detector. Every (model, detector) pair draws from its own random generator, derived from the top-level seed and the detector name via an independent SeedSequence stream.

A model carrying a :class:~gwmock_noise.glitches.models.NetworkCoherence runs one Poisson process for the whole network and draws one waveform per event, offering it to every detector it applies to. GlitchModel.rate is then the network event rate, not the per-detector one, and each detector sees it times the coherence's participation probability. Arrival times and waveforms come from a stream keyed by the model alone; each detector's participation and amplitude ratio come from a stream keyed by its own name and the event's ordinal.

Either way, a detector's realization is reproducible and independent of which other detectors are present or the order in which they are requested -- which is what lets two runs over different networks be compared on the channels they share.

A glitch whose waveform extends past the end of a chunk has its remainder carried into the next chunk and replayed in event order, so the injected glitch series is identical, sample for sample, to a single generate call of the same total duration (the per-sample addition order is preserved exactly, not merely up to rounding). A tail is dropped where the data it would continue into does not exist: past the last generated chunk, and past a segment whose successor the caller places at a discontinuous epoch.

Every event that lands in the strain is recorded as it fires, not merely tallied: see :attr:glitch_events for the cumulative truth catalogue and :attr:segment_glitch_events for the events the most recent generate call injected. A row is written where a glitch starts, so a waveform straddling a chunk boundary appears once, in the chunk that holds its first sample, at its true time -- the carried-over tail adds samples to the next chunk but no second row. GLITCH_CATALOGUE_COLUMNS documents the columns and GLITCH_CATALOGUE_TIME_CONVENTION which instant each time column names.

gps_start is the GPS time of the first sample of the next segment, and it belongs to the caller. It advances by each generated segment's duration, so a stream's catalogue reads in real GPS time on its own; a caller that knows better -- a writer placing non-contiguous segments, or one driving the stream batch by batch -- assigns it before each generate call and that assignment is honoured, including on a call that also carries a seed. Only reset, the explicit start-over, rewinds it to the epoch the wrapper was constructed with.

glitch_events property

Return every glitch injected since the last (re)seed, in time order.

Copies, so a consumer cannot edit the run's own record of itself.

metadata property

Return additive-wrapper metadata.

segment_glitch_events property

Return the glitches the most recent generate call injected.

What a per-batch consumer wants: the rows belonging to the chunk it is about to write, rather than the whole stream's catalogue re-filtered by time. Empty before the first generate call, and after a call that fired no events.

__init__(base, glitch_models, gps_start=0.0)

Initialize the additive glitch wrapper.

generate(duration, sampling_frequency, detectors, seed=None)

Generate base noise and inject the configured glitches.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield glitch-injected chunks lazily.

reset()

Reset the additive wrapper and any resettable base state.

Rewinds the epoch to the one the wrapper was constructed with, which _initialize_process does not: this is the caller saying "start over", where a seed passed to generate is the caller saying "this segment, whose epoch I have just told you, begins a new realization". The two need different answers, and conflating them is what put every streamed segment on the wrong epoch.

JointCovariance dataclass

A one-sided cross-spectral density matrix over a set of channels.

The convention matches :func:numpy.fft.rfftfreq: frequencies is the one-sided, non-negative frequency grid, and matrix[k] is the Hermitian (n_channels, n_channels) one-sided cross-spectral density at frequencies[k], in units of channel_unit**2 / Hz. The diagonal entry matrix[k, i, i] is the one-sided PSD of channel channel_order[i] at that frequency.

Attributes:

Name Type Description
frequencies ndarray

One-sided frequency grid in hertz, shape (n_freq,).

matrix ndarray

Cross-spectral density matrices, shape (n_freq, n, n), complex-valued and Hermitian at every frequency.

channel_order list[str]

Channel names, in the order matrix's last two axes index them.

JointDummyCorrelatedSimulator

Public dummy backend producing simultaneous, correlated strain and witness channels.

Every channel is temporally white with unit per-sample variance and coupling correlation with every other channel; there is no spectral shaping and no physical model. The one-sided cross-spectral density is therefore flat and known in closed form: for per-sample covariance Sigma at sampling frequency fs, the one-sided PSD/CSD is the frequency-independent matrix 2 * Sigma / fs (the same covariance-to-PSD scale convention already used by :class:~gwmock_noise.simulators.multichannel.MultichannelNoiseSimulator, inverted).

channel_metadata property

Return typed metadata for every strain and witness channel.

metadata property

Return simulator metadata.

__init__(*, detectors, witnesses, sampling_frequency=DEFAULT_SAMPLING_FREQUENCY, duration=DEFAULT_DURATION, seed=None, coupling=DEFAULT_COUPLING, strain_unit='strain', witness_unit='counts')

Initialize the dummy backend.

Parameters:

Name Type Description Default
detectors list[str]

Strain channel names. Fixed for the lifetime of this instance; later calls may reorder but not change this set.

required
witnesses list[str]

Witness channel names, disjoint from detectors. Fixed for the lifetime of this instance in the same way.

required
sampling_frequency float

Sampling frequency in hertz.

DEFAULT_SAMPLING_FREQUENCY
duration float

Nominal generation duration in seconds.

DEFAULT_DURATION
seed int | None

Seed for the shared innovation generator.

None
coupling float

Equicorrelation coefficient applied between every pair of channels (strain-strain, witness-witness and strain-witness alike). Must satisfy 0.0 <= coupling < 1.0.

DEFAULT_COUPLING
strain_unit str

Unit string recorded for strain channels.

'strain'
witness_unit str

Unit string recorded for witness channels.

'counts'

Raises:

Type Description
ValueError

If detectors/witnesses are empty, contain duplicates, overlap each other, or if coupling is out of range.

covariance()

Return this backend's exact, closed-form one-sided cross-spectral density.

Every channel is temporally white, so the returned matrix is frequency-independent: 2 * Sigma / sampling_frequency at every frequency, where Sigma is the per-sample covariance matrix drawn from at construction. This is an exact analytic value, not a statistical estimate off a realization.

generate_joint(duration, sampling_frequency, detectors, witnesses, seed=None)

Generate one simultaneous, correlated strain+witness realization.

generate_joint_stream(chunk_duration, sampling_frequency, detectors, witnesses, seed=None)

Yield simultaneous strain+witness chunks lazily, preserving RNG state.

JointRealization dataclass

One simultaneous draw of strain and witness channels.

Attributes:

Name Type Description
strain dict[str, ndarray]

Per-detector strain arrays, one entry per requested detector.

witness dict[str, ndarray]

Per-channel witness arrays, one entry per requested witness channel.

channel_metadata dict[str, ChannelMetadata]

Typed metadata for every channel in strain and witness, keyed by the same channel names.

provenance dict[str, Any]

Implementation-defined provenance describing how this realization was produced (backend identity, package version, seed, RNG algorithm, and any upstream configuration actually used).

JointStrainWitnessSimulator

Bases: Protocol

Structural interface for simulators producing joint strain+witness output.

This is an optional extension of the legacy NoiseSimulator contract: implementing it changes nothing about NoiseSimulator itself, and a NoiseSimulator-only consumer is unaffected by its existence. A class may satisfy both protocols at once.

metadata property

Return simulator metadata (same role as NoiseSimulator.metadata).

covariance()

Return the joint one-sided cross-spectral density over all channels.

generate_joint(duration, sampling_frequency, detectors, witnesses, seed=None)

Generate one simultaneous strain+witness realization.

generate_joint_stream(chunk_duration, sampling_frequency, detectors, witnesses, seed=None)

Yield simultaneous strain+witness chunks lazily.

Continuity contract: consecutive chunks from this iterator must equal the same realization a caller would obtain from one seeded :meth:generate_joint call spanning the combined duration -- the same contract NoiseSimulator.generate_stream makes for strain-only output.

LogNormalAmplitudeDistribution dataclass

Log-normal amplitude sampler parameterized by linear mean and std.

__post_init__()

Validate the configured distribution parameters.

sample(rng)

Draw one amplitude sample.

MultichannelNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate multichannel noise from a tabulated PSD/CSD matrix.

ar_coefficients property

Return a copy of the fitted coefficient matrices A_1 .. A_p.

design_frequencies property

Return a copy of the design-grid frequencies in hertz.

innovation_covariance property

Return a copy of the fitted innovation covariance.

metadata property

Return metadata describing the fitted matrix spectral factor.

model_spectral_matrices property

Return a copy of the fitted cross-spectral matrices on the design grid.

model_spectral_matrix_curve property

Return the band-masked fitted cross-spectral matrices on the design grid.

resume_metadata_nbytes property

Serialized size of the resume metadata in bytes.

state_nbytes property

Serialized size of the continuation state in bytes.

target_spectral_matrices property

Return a copy of the target cross-spectral matrices on the design grid.

The matrices are in the covariance scale the recursion uses (one_sided_psd * sampling_frequency / 2) and carry the band mask and edge taper the fit was built from.

target_spectral_matrix_curve property

Return the band-masked target cross-spectral matrices on the design grid.

__init__(*, psd_files=None, csd_files=None, target_matrices=None, target_frequencies=None, order=DEFAULT_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization_epsilon=DEFAULT_REGULARIZATION_EPSILON)

Initialize the simulator and fit the matrix spectral factor once.

The cross-spectral matrix may be given either as files (psd_files plus csd_files, absent off-diagonal pairs mean zero coherence) or as an in-memory target_matrices array on target_frequencies, which lets an analytic pair bypass the tabulated-curve interpolation.

Parameters:

Name Type Description Default
psd_files dict[str, str | Path] | None

Per-detector PSD table paths, keyed by detector.

None
csd_files dict[str, str | Path] | dict[tuple[str, str], str | Path] | None

Cross-spectral table paths, keyed by detector pair or by "DET1-DET2" strings.

None
target_matrices ndarray | None

One-sided cross-spectral matrices of shape (n_frequencies, n_channels, n_channels), mutually exclusive with the file inputs.

None
target_frequencies ndarray | None

Frequencies of target_matrices in hertz.

None
order int

Order p of the autoregressive factor.

DEFAULT_ORDER
detectors list[str] | None

Detector names to generate.

None
sampling_frequency float

Sampling frequency in hertz.

4096.0
duration float

Nominal generation duration in seconds.

4.0
seed int | None

Seed for the shared innovation generator.

None
low_frequency_cutoff float

Lower edge of the fit band in hertz.

2.0
high_frequency_cutoff float | None

Upper edge of the fit band in hertz.

None
block_size int

Number of samples generated per recursion call.

DEFAULT_BLOCK_SIZE
regularization_epsilon float

Relative ridge, as a fraction of the largest in-band diagonal entry, added as a zero-lag (white) floor to the autocovariance when the in-band target is not positive definite. Zero forbids the ridge and makes such a target a fit failure.

DEFAULT_REGULARIZATION_EPSILON

continuation_state()

Return the bounded state needed to keep a running stream going.

The state is the recursion history plus the pending output; its size is proportional to the order and independent of the generated span.

export_state()

Return a picklable snapshot sufficient to resume the stream.

from_component(component, config) classmethod

Construct a multichannel simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector multichannel noise with continuity across calls.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield multichannel chunks lazily while preserving recursion history.

import_state(state)

Restore a snapshot produced by :meth:export_state.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same configuration.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator's settings.

reset()

Clear the recursion history, pending output and generator.

resume_metadata()

Return the metadata a stopped stream must persist to resume.

NoiseSimulator

Bases: Protocol

Structural interface for downstream noise simulators.

metadata property

Return simulator metadata.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector strain arrays.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield per-detector strain chunks lazily.

OverlapSaveFirSimulator

Bases: ConfigurableNoiseSimulator

Generate coloured detector noise with a minimum-phase overlap-save FIR.

filter_taps property

Return a copy of the designed colouring filter taps.

metadata property

Return simulator metadata.

resume_metadata_nbytes property

Serialized size of the resume metadata in bytes.

state_nbytes property

Serialized size of the continuation state in bytes.

target_psd property

Return a copy of the band-masked target PSD the filter was designed from.

__init__(*, psd_file=None, target_psd=None, target_frequencies=None, filter_length=DEFAULT_FILTER_LENGTH, minimum_phase=True, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE)

Initialize the simulator and design the colouring filter once.

The target may be given either as a file (psd_file) or directly as a one-sided array (target_psd), so an analytic target can bypass the tabulated-curve interpolation entirely. When target_frequencies is omitted for an array target it defaults to the one-sided grid matching its length at the configured sampling frequency.

Parameters:

Name Type Description Default
psd_file str | Path | None

Path to a two-column PSD table.

None
target_psd ndarray | None

One-sided target PSD array, mutually exclusive with psd_file.

None
target_frequencies ndarray | None

Frequencies of target_psd in hertz.

None
filter_length int

Truncation control L_f in samples.

DEFAULT_FILTER_LENGTH
minimum_phase bool

Whether to apply the cepstral minimum-phase factorisation.

True
detectors list[str] | None

Detector names to generate.

None
sampling_frequency float

Sampling frequency in hertz.

4096.0
duration float

Nominal generation duration in seconds.

4.0
seed int | None

Base seed for the per-detector generators.

None
low_frequency_cutoff float

Lower edge of the target band in hertz.

2.0
high_frequency_cutoff float | None

Upper edge of the target band in hertz.

None
block_size int

Overlap-save block length in samples.

DEFAULT_BLOCK_SIZE

continuation_state()

Return the bounded state needed to keep a running stream going.

The state is the filter memory of every detector plus the pending output and block bookkeeping. It contains no generated strain and its size is independent of the generated span.

export_state()

Return a picklable snapshot sufficient to resume the stream.

from_component(component, config) classmethod

Construct an overlap-save FIR simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector coloured noise with continuity across calls.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield coloured-noise chunks lazily while preserving filter memory.

import_state(state)

Restore a snapshot produced by :meth:export_state.

Parameters:

Name Type Description Default
state dict[str, Any]

A snapshot from another simulator with the same configuration.

required

Raises:

Type Description
ValueError

If the snapshot does not match this simulator's settings.

reset()

Clear the filter memory, pending samples and random-number generators.

resume_metadata()

Return the metadata a stopped stream must persist to resume.

This is the per-detector bit-generator state plus the settings needed to rebuild the filter, and is reported separately from the continuation state.

ScatteredLightGlitch dataclass

Bases: GlitchModel

Arch-shaped scattered-light transient with a Gaussian envelope.

__post_init__()

Validate scattered-light parameters.

applies_to(detector)

Return whether this model injects glitches into detector.

Parameters:

Name Type Description Default
detector str

The interferometer name.

required

Returns:

Type Description
bool

True when the model has no selector, or names this interferometer.

coloring_reference()

Return a stable identity for the PSD this model colors against.

None for a model that does no coloring, which therefore imposes no noise floor on the interferometers it claims and cannot contradict another model about one. Read from psd_file where the model has one, so a third-party model gets the same treatment as the built-in ones without restating the attribute name.

Returns:

Type Description
str | None

The resolved PSD identity, or None when the model is uncolored.

draw(sampling_frequency, rng=None)

Generate one waveform together with the parameters that produced it.

What the injector calls, so every event can be recorded rather than merely tallied.

generate_waveform remains the extension point it has always been: a model that overrides it -- including a subclass of a built-in model -- governs what is injected, and its draw is reported with the parameters blank rather than filled in from the superclass's draw, which produced a different waveform. Reversing that precedence would let an override change the strain while the catalogue described the code it replaced.

generate_waveform(sampling_frequency, rng=None)

Generate a chirping scattered-light glitch.

resolve()

Pin any external, mutable dependency to an immutable version.

Parametric models are fully specified by their configuration and have nothing external to resolve, so the base implementation is a no-op that returns None. Models backed by a downloaded dataset (e.g. DeepExtractorGlitch) override this to fetch and return the concrete version their run is pinned to. Exposing it on the base lets a caller pin every model in a heterogeneous glitch list uniformly.

serialize()

Return metadata-friendly model parameters.

SchumannNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate correlated strain noise from an isotropic Schumann-resonance model.

metadata property

Return simulator metadata.

previous_strain property

Expose continuity buffers for protocol-compatible state inspection.

__init__(*, positions, coupling_files, detectors=None, schumann_params=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, window_duration=DEFAULT_WINDOW_DURATION)

Initialize the Schumann simulator.

from_component(component, config) classmethod

Construct a Schumann-noise simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate correlated Schumann strain noise.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield Schumann-noise chunks lazily while preserving simulator state.

reset()

Clear continuity and RNG state.

spectrum(frequencies)

Return the isotropic Schumann magnetic PSD on a frequency grid.

theoretical_coherence(frequency, detector_a, detector_b)

Return the isotropic Schumann coherence approximation.

SchumannParams dataclass

Physical parameters for the isotropic Schumann-resonance model.

__post_init__()

Validate the resonance parameter vectors.

SimulationResult dataclass

Result of a noise simulation run.

Attributes:

Name Type Description
output_paths dict[str, Path]

Paths to generated output files, keyed by detector name.

config NoiseConfig

The configuration used for the simulation.

SpectralLineSimulator

Bases: ConfigurableNoiseSimulator

Generate additive spectral lines directly in the time domain.

metadata property

Return simulator metadata.

__init__(*, lines, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None)

Initialize the line generator.

from_component(component, config) classmethod

Construct a spectral-line simulator from one component definition.

generate(duration, sampling_frequency, detectors, seed=None)

Generate line-only strain for all requested detectors.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield spectral-line chunks lazily.

reset()

Clear continuity state and phase initialization.

WhiteNoiseSimulator

Bases: ConfigurableNoiseSimulator

Generate Gaussian white noise as a composable component.

metadata property

Return metadata describing the white-noise component.

__init__(*, duration=4.0, sampling_frequency=4096.0, detectors=None, seed=None)

Initialize the white-noise component state.

from_component(component, config) classmethod

Construct one white-noise component from generic config.

generate(duration, sampling_frequency, detectors, seed=None)

Return Gaussian white-noise strain arrays.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield white-noise strain chunks lazily.

apply_segment_gps_start(simulator, gps_start)

Point every glitch injector inside a simulator chain at one segment epoch.

A writer that assigns a per-segment epoch -- FrameWriter, whose frame names carry it -- is the authority on where a segment sits in GPS time, and the glitch truth catalogue has to be stamped with that same instant or the catalogue and the frame name describe different data. Threading the writer's epoch through is what keeps it one number rather than two that agree only while the segments happen to be contiguous.

Walks the wrapper chain (base) and a composite's components, because the injector is rarely the outermost simulator: a glitch component sits inside CompositeNoiseSimulator, which sits inside whatever adapter the writer holds. A simulator with no injector anywhere beneath it is left alone.

A wrapper that builds its simulators rather than holding one -- ParallelAdapter, which constructs and caches one per detector from a factory -- cannot be reached by walking attributes, and an epoch set on such a wrapper alone reached none of its workers. Those declare a writable segment_gps_start and own the fan-out, so this hands the epoch over and lets them place both the workers they already hold and any they build later in the same segment.

available_joint_backend_names()

Return the registered joint backend names, sorted.

discover_joint_backends() cached

Return the entry points registered under :data:JOINT_BACKEND_ENTRY_POINT_GROUP.

This performs discovery only -- it does not import any backend's module. Call :meth:importlib.metadata.EntryPoint.load on a returned entry point (or use :func:load_joint_backend) to import and resolve the backend class it names.

load_joint_backend(name)

Import and return the backend class registered under name.

Parameters:

Name Type Description Default
name str

The entry-point name a package registered its backend under.

required

Returns:

Type Description
Any

The backend class (or factory callable) the entry point names.

Raises:

Type Description
KeyError

If no backend is registered under name.

open_stream(simulator, *, chunk_duration, sampling_frequency, detectors, seed=None)

Open a stateful chunk stream on a protocol-compatible simulator.

take(stream, total_duration, chunk_duration, sampling_frequency)

Collect stream chunks up to total_duration seconds.

whittle_levinson_factorization(autocovariance, order)

Factorize a matrix autocovariance with the Whittle block recursion.

Parameters:

Name Type Description Default
autocovariance ndarray

The autocovariance matrices R_0 .. R_p, shape (order + 1, n_channels, n_channels). The block-Toeplitz matrix they build must be positive definite.

required
order int

Model order p (at least one).

required

Returns:

Type Description
WhittleFactorization

The fitted autoregressive coefficients, innovation covariance and the

WhittleFactorization

conditioning diagnostics.

Raises:

Type Description
FitError

If the sequence is too short, not finite, or not positive definite at some order of the recursion.

ValueError

If order is not a positive integer.

Joint-backend conformance suite

The shared contract tests a JointStrainWitnessSimulator backend runs under its own package's test runner (requires pytest).

A shared conformance suite for :class:JointStrainWitnessSimulator backends.

Every package that registers a joint strain+witness backend can run exactly the same contract checks against it, under its own test runner. gwmock-noise runs this suite against its public dummy backend; a downstream package runs it against its own backends.

Usage, in a consumer's test module::

import pytest
from gwmock_noise import load_joint_backend
from gwmock_noise.testing.joint_conformance import JointBackendCase, JointBackendConformance


def _build(seed):
    return load_joint_backend("my_backend")(sampling_frequency=64.0, seed=seed)


@pytest.fixture(params=[JointBackendCase("my_backend", _build, 64.0, "my_backend")], ids=lambda case: case.name)
def joint_backend(request):
    return request.param


class TestMyBackendConformance(JointBackendConformance):
    pass

The suite checks the realization and metadata shapes, seeded determinism, stream continuity, a Hermitian positive-semidefinite covariance over exactly the realized channels, JSON-native provenance and metadata, refusal of an unknown channel, and that no witness is disguised as a strain channel or carried as an array in metadata.

This module requires pytest.

JointBackendCase dataclass

One joint backend under test, and how to build it.

Attributes:

Name Type Description
name str

Label for the case, used in test ids.

build Callable[[int | None], JointStrainWitnessSimulator]

Builds a fresh simulator; called with the construction seed, or None when the seed is passed to the generation call instead.

sampling_frequency float

The sampling frequency the built simulator runs at.

provenance_backend str

The value provenance["backend"] must carry.

chunk_samples int

Samples per generated chunk.

chunk_duration property

Duration of one chunk in seconds.

generate(simulator, seed=None, chunks=1)

Generate chunks chunks' worth of the simulator's own channels in one call.

JointBackendConformance

The contract every joint backend must meet.

Subclass it under a Test-prefixed name and provide a joint_backend fixture yielding :class:JointBackendCase instances.

Beyond the structural protocol, the suite holds a backend to these conventions: provenance records backend, package_version, seed and rng_bit_generator; metadata lists the channel names under detectors and witnesses; and a channel name the simulator was not built with raises a :class:ValueError whose message says the change is unsupported.

test_an_unknown_channel_is_refused(joint_backend)

An unknown detector or witness name is refused rather than silently generated.

test_channel_metadata_is_typed_and_complete(joint_backend)

Every realized channel has typed metadata with the right kind, domain, rate and dtype.

test_covariance_is_hermitian_psd_over_the_realized_channels(joint_backend)

The covariance indexes exactly the realized channels and is Hermitian PSD on an rfft grid.

test_no_channel_data_in_metadata_or_provenance(joint_backend)

Metadata, provenance and channel metadata hold no arrays and no channel-length sequences.

test_provenance_and_metadata_are_json_native(joint_backend)

Provenance and metadata survive a strict JSON round trip and record backend, version, seed and RNG.

test_realization_shape_dtype_and_channel_sets(joint_backend)

Strain and witness dicts hold exactly the requested channels as finite 1-D float64 arrays.

test_satisfies_the_public_protocol(joint_backend)

The runtime-checkable public protocol recognizes the backend.

test_seeded_generation_is_deterministic(joint_backend)

Equal seeds give bit-identical realizations; a different seed changes the strain.

test_stream_chunks_continue_the_batch(joint_backend)

Four seeded stream chunks concatenate to the seeded batch over the same span.

test_witnesses_are_never_strain_channels(joint_backend)

Detector and witness names are disjoint, and every witness is typed and delivered as a witness.

array_leaks(value, path='')

Return the paths in a metadata/provenance tree that hold channel-like data.

Flags any :class:numpy.ndarray, and any list or tuple of more than :data:MAX_METADATA_VECTOR numbers. Sequences of names (channel orders) are not data and are walked, not flagged.

Parameters:

Name Type Description Default
value Any

The tree to walk.

required
path str

The path of value within the enclosing tree.

''

Returns:

Type Description
list[str]

One description per leaking path; empty when nothing leaks.

Parallel execution

ParallelAdapter runs independent-detector simulators across threads or processes (with constraints for correlated backends).

Parallel adapter for independent-detector noise simulators.

ParallelAdapter

Wrap an independent-detector simulator factory with parallel execution.

metadata property

Return metadata describing the wrapped simulator and executor.

segment_gps_start property writable

Return the GPS epoch assigned for the next segment, or None if never set.

__init__(base_factory, *, max_workers=None, backend='auto')

Store the simulator factory and executor preferences.

generate(duration, sampling_frequency, detectors, seed=None)

Generate per-detector strain arrays in parallel.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield parallel-generated chunks lazily.

reset()

Clear cached worker simulators, resetting continuity across workers.

Diagnostics

PSD estimation/comparison and simple stationarity/gaussianity checks used in tests and notebooks.

Diagnostics helpers for gwmock-noise.

DiagnosticResult dataclass

Outcome of a statistical diagnostic check.

compare_psd(data, target_psd_file, sampling_frequency, rtol=0.1, fmin=10.0, fmax=None)

Compare an estimated PSD against a target PSD file over a frequency band.

This helper compares the median PSD level in-band rather than every bin individually, which makes it robust to the residual variance of a single Welch estimate from stochastic noise.

estimate_psd(data, sampling_frequency, segment_duration=4.0, overlap=0.5)

Estimate a one-sided PSD in physical density units.

The returned frequencies are in Hz. For strain input data, the PSD is in strain^2/Hz.

plot_psd(frequencies, psd, ax=None)

Plot a PSD on log-log axes, importing matplotlib only when needed.

run_diagnostics(data, sampling_frequency, whiten=True)

Run the default Gaussianity and stationarity diagnostics.

Data is whitened against its own Welch PSD estimate before testing (see _whiten_samples), which makes both checks applicable to coloured noise. Pass whiten=False to test the raw samples instead — only meaningful for broadband data.

Output adapters

Lazy exports for GWpyAdapter and FrameWriter (optional gwpy / frame extras).

Output adapters for gwmock-noise.

__getattr__(name)

Lazily resolve optional adapter exports.

GWOSC real-noise fetching

Models, segment filtering, data fetching, and a NoiseSimulator wrapper for retrieving real detector strain data from GWOSC.

GWOSC real-noise fetching subpackage.

Provides configuration models, segment filtering, and data fetching for retrieving real gravitational-wave detector strain data from the Gravitational-Wave Open Science Centre (GWOSC).

FilterType

Bases: StrEnum

Types of segments that can be filtered out from GWOSC data.

GwoscFilterConfig

Bases: BaseModel

Configuration for filtering GWOSC strain segments.

Attributes:

Name Type Description
filter_types list[FilterType]

Which categories of segments to exclude.

far_threshold float

Maximum false-alarm rate (events/year) for GW event filtering. Events with FAR above this threshold are excluded from filtering (treated as noise). Default 1.0 (one per year).

event_padding float

Padding in seconds to apply around each GW event GPS time when building vetosegments.

dq_flags list[str]

DQ flag basenames (without detector prefix) to query for data-quality vetosegments. E.g. ["CBC_CAT1", "CBC_CAT2"]. The detector prefix is prepended automatically.

exclude_hardware_injections bool

Whether to also exclude segments containing hardware injections.

has_dq_filter property

Return whether the data-quality filter is active.

has_gw_filter property

Return whether any GW signal filter is active.

include_marginal_events property

Return whether to include marginal (low-confidence) GW events.

GwoscNoiseConfig

Bases: BaseModel

Configuration for fetching real detector noise from GWOSC.

Attributes:

Name Type Description
detectors list[str]

List of detector prefixes (e.g. ["H1", "L1"]).

gps_start float

GPS start time of the requested data interval.

gps_end float

GPS end time of the requested data interval.

sample_rate float

Desired sampling rate in Hz. GWOSC typically provides 4096 Hz, with some event datasets at 16384 Hz.

filters GwoscFilterConfig

Filter configuration for excluding segments.

host str

GWOSC host URL.

cache_dir Path | None

Optional directory for file-level caching. When set, downloaded HDF5 frame files are saved locally and reused on subsequent requests for the same GPS interval. When None, data is downloaded on every call without caching.

duration property

Return the duration of the requested interval in seconds.

validate_gps_range()

Validate that gps_end is after gps_start.

GwoscNoiseFetcher

Fetch real detector noise data from GWOSC with optional filtering.

Downloads strain data from GWOSC and applies user-configured filters to exclude segments containing GW signals and data-quality issues.

When cache_dir is configured, HDF5 files are saved locally and reused on subsequent requests — avoiding repeated downloads for the same GPS interval.

Attributes:

Name Type Description
config

The GWOSC noise fetching configuration.

clean_segments property

Return the computed clean segments per detector.

Returns:

Type Description
dict[str, list[tuple[float, float]]]

A dictionary mapping each detector to a list of

dict[str, list[tuple[float, float]]]

(start, end) clean segment tuples.

__init__(config)

Initialize the fetcher.

Parameters:

Name Type Description Default
config GwoscNoiseConfig

Configuration specifying detectors, GPS range, sample rate, filtering options, and optional cache directory.

required

check_availability(detectors=None)

Probe GWOSC for per-detector strain-data availability.

Unlike :attr:clean_segments, which only reflects vetoes (GW events and data-quality flags), this checks whether the strain data itself has been published and is downloadable for the full configured GPS interval. A detector can have a fully "clean" interval yet have no published strain — e.g. the data has not been released yet for the observing run — in which case a fetch would fail. Use this for a pre-flight check before fetching.

Parameters:

Name Type Description Default
detectors list[str] | None

Detectors to probe. Defaults to all configured detectors when None.

None

Returns:

Type Description
dict[str, bool]

A dictionary mapping each requested detector to True if

dict[str, bool]

GWOSC has data URLs covering the interval, False otherwise.

fetch_clean(detectors=None)

Fetch clean noise segments.

Clean segments are computed by excluding GW events and data-quality issues according to the filter configuration.

Parameters:

Name Type Description Default
detectors list[str] | None

Detectors to fetch. Defaults to all configured detectors when None.

None

Returns:

Type Description
dict[str, list[TimeSeries]]

A dictionary mapping each requested detector to a list of

dict[str, list[TimeSeries]]

gwpy.TimeSeries, one per clean segment.

Raises:

Type Description
ValueError

If no data is available for any requested detector or no clean segments are found.

fetch_raw(detectors=None)

Fetch raw strain data without filtering.

Parameters:

Name Type Description Default
detectors list[str] | None

Detectors to fetch. Defaults to all configured detectors when None.

None

Returns:

Type Description
dict[str, TimeSeries]

A dictionary mapping each requested detector to a full-interval

dict[str, TimeSeries]

gwpy.TimeSeries.

Raises:

Type Description
ValueError

If no data is available for any requested detector.

GwoscSegmentFilter

Compute clean analysis segments from GWOSC metadata.

Queries GWTC event catalogs and data-quality flags to build vetosegments (time windows to exclude) and returns the remaining clean intervals suitable for noise analysis.

Attributes:

Name Type Description
config

The filter configuration.

__init__(config)

Initialize the segment filter.

Parameters:

Name Type Description Default
config GwoscFilterConfig

Filter configuration specifying which segment categories to exclude and their parameters.

required

compute_clean_segments(gps_start, gps_end, detectors)

Compute clean (analysis-ready) segments for each detector.

Clean segments are the requested GPS range minus the union of: - GW event vetosegments (if any GW filter is active) - DQ vetosegments (if the data-quality filter is active) - Hardware injection segments (if enabled)

Parameters:

Name Type Description Default
gps_start float

GPS start of the requested interval.

required
gps_end float

GPS end of the requested interval.

required
detectors list[str]

List of detector prefixes.

required

Returns:

Type Description
dict[str, SegmentList]

A dictionary mapping each detector to a list of clean

dict[str, SegmentList]

(start, end) segments.

get_dq_vetosegments(gps_start, gps_end, detector)

Query data-quality flags for detector and build vetosegments.

Uses gwosc.timeline.get_segments() to fetch pre-computed DQ veto segments for each flag in dq_flags.

Parameters:

Name Type Description Default
gps_start float

GPS start of the query interval.

required
gps_end float

GPS end of the query interval.

required
detector str

Detector prefix (e.g. "H1").

required

Returns:

Type Description
SegmentList

A list of (start, end) vetosegment tuples covering

SegmentList

time windows with data-quality issues.

get_gw_vetosegments(gps_start, gps_end)

Query GWTC events in the GPS range and build vetosegments.

For HIGH_CONFIDENCE_GW, only events with FAR <= far_threshold are included. For ALL_GW_SIGNALS, all events in the range are included.

Parameters:

Name Type Description Default
gps_start float

GPS start of the query interval.

required
gps_end float

GPS end of the query interval.

required

Returns:

Type Description
SegmentList

A list of (start, end) vetosegment tuples, each centred

SegmentList

on a GW event GPS time with event_padding on both sides.

get_hardware_injection_segments(gps_start, gps_end)

Query hardware injection events and build vetosegments.

Parameters:

Name Type Description Default
gps_start float

GPS start of the query interval.

required
gps_end float

GPS end of the query interval.

required

Returns:

Type Description
SegmentList

A list of (start, end) vetosegment tuples around

SegmentList

hardware injection times.

Real-noise simulator backed by GWOSC strain data.

Provides a :class:GwoscNoiseSimulator that satisfies the NoiseSimulator protocol, allowing GWOSC-fetched noise to be used interchangeably with the built-in synthetic simulators.

GwoscNoiseSimulator

Real-noise simulator that fetches strain data from GWOSC.

Implements the NoiseSimulator protocol so it can be used interchangeably with synthetic simulators like ColoredNoiseSimulator and CorrelatedNoiseSimulator.

Fetches real detector strain from the Gravitational-Wave Open Science Centre and applies user-configured filters to exclude segments containing GW signals or data-quality issues.

When cache_dir is set in the config, downloaded HDF5 files are saved locally and reused — avoiding repeated downloads for the same GPS interval.

Attributes:

Name Type Description
duration float

Duration of the configured GPS interval (seconds).

sampling_frequency float

Sampling frequency in Hz.

detectors list[str]

List of detector prefixes.

seed None

Always None (real noise has no random seed).

config

The underlying GWOSC configuration.

detectors property

Return the list of detector prefixes.

duration property

Return the total duration of the configured GPS interval.

metadata property

Return metadata describing the simulator and its configuration.

Returns:

Type Description
dict[str, Any]

A dictionary with the simulator implementation name,

dict[str, Any]

GPS range, sample rate, detectors, filter configuration,

dict[str, Any]

and cache status.

sampling_frequency property

Return the sampling frequency in Hz.

seed property

Return None — real noise has no controllable random seed.

__init__(config)

Initialize the real-noise simulator.

Parameters:

Name Type Description Default
config GwoscNoiseConfig

Configuration specifying GPS range, detectors, sample rate, filtering options, and optional cache directory.

required

check_availability(detectors=None)

Probe GWOSC for per-detector strain-data availability.

Convenience passthrough to :meth:GwoscNoiseFetcher.check_availability. Use this as a pre-flight check: a detector may have a fully clean interval (no vetoes) yet have no published strain data, in which case :meth:generate would fail with a clear error.

Parameters:

Name Type Description Default
detectors list[str] | None

Detectors to probe. Defaults to all configured detectors when None.

None

Returns:

Type Description
dict[str, bool]

A dictionary mapping each requested detector to True if

dict[str, bool]

GWOSC has data covering the interval, False otherwise.

generate(duration, sampling_frequency, detectors, seed=None)

Fetch clean noise and return the longest contiguous segment per detector.

Real GWOSC noise is fragmented by the excision of GW events and data-quality vetoes, so the clean data generally consists of several non-contiguous spans. To honour the NoiseSimulator contract — a single, gap-free array per detector — this returns the longest contiguous clean segment for each detector.

Concatenating the fragments instead (the historical behaviour) splices non-adjacent samples together, creating discontinuities that show up as sharp transients after whitening. Use :meth:generate_segments to obtain every clean segment separately when the full clean data is required.

Parameters:

Name Type Description Default
duration float

Requested duration (ignored; the longest available contiguous clean segment determines the output length).

required
sampling_frequency float

Requested sampling frequency (must match the configured sample_rate).

required
detectors list[str]

Requested detector list (must be a subset of the configured detectors).

required
seed int | None

Ignored for real noise.

None

Returns:

Type Description
dict[str, ndarray]

A dictionary mapping each detector to a 1-D numpy array

dict[str, ndarray]

holding its longest contiguous clean segment.

Raises:

Type Description
ValueError

If sampling_frequency does not match the configured value, if detectors are not a subset, or if a detector has no clean data.

generate_segments(duration, sampling_frequency, detectors, seed=None)

Fetch clean noise and return per-detector lists of contiguous segments.

Unlike :meth:generate, this preserves the gap structure of the real data: each clean segment — the spans left after excluding GW events and data-quality vetoes — is returned as its own array, in GPS-time order. Adjacent segments are not contiguous in time, so callers must treat each segment independently; concatenating them would splice non-adjacent samples together and introduce discontinuities at the joins (which manifest as sharp transients after whitening).

Parameters:

Name Type Description Default
duration float

Requested duration (ignored; the GPS interval from the config determines the available data).

required
sampling_frequency float

Requested sampling frequency (must match the configured sample_rate).

required
detectors list[str]

Requested detector list (must be a subset of the configured detectors).

required
seed int | None

Ignored for real noise.

None

Returns:

Type Description
dict[str, list[ndarray]]

A dictionary mapping each detector to a list of 1-D numpy

dict[str, list[ndarray]]

arrays, one per clean segment, ordered by GPS time.

Raises:

Type Description
ValueError

If sampling_frequency does not match the configured value, if detectors are not a subset, or if a detector has no clean data.

generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)

Yield clean-noise chunks lazily.

Fetches the longest contiguous clean segment once (see :meth:generate) and yields it in chunks of chunk_duration seconds. Chunking the longest contiguous span keeps every chunk free of the discontinuities that splicing fragments would introduce.

Parameters:

Name Type Description Default
chunk_duration float

Duration of each yielded chunk in seconds.

required
sampling_frequency float

Requested sampling frequency (must match the configured sample_rate).

required
detectors list[str]

Requested detector list (must be a subset of the configured detectors).

required
seed int | None

Ignored for real noise.

None

Yields:

Type Description
dict[str, ndarray]

Per-detector strain arrays for each chunk.

Command-line interface

Typer application and the simulate command used by the gwmock-noise console script.

Main entry point for the gwmock_noise CLI application.

LoggingLevel

Bases: StrEnum

Logging levels for the CLI.

main(verbose=LoggingLevel.INFO)

Implement the main entry point for the CLI application.

Parameters:

Name Type Description Default
verbose Annotated[LoggingLevel, Option('--verbose', '-v', help='Set verbosity level.')]

Verbosity level for logging.

INFO

register_commands()

Register CLI commands.

setup_logging(level=LoggingLevel.INFO)

Set up logging with Rich handler.

Parameters:

Name Type Description Default
level LoggingLevel

Logging level.

INFO

Command for running noise simulations.

simulate(config_path)

Run noise simulation from a configuration file.

Loads the configuration, validates it, and runs the noise simulator. Output files are written to the directory specified in the config.

Package version

A script to infer the version number from the metadata.