Overview
Reference material is generated from Python docstrings via mkdocstrings. The sections below mirror the main import paths used in applications and in the CLI.
Top-level package¶
Re-exports for common workflows (NoiseConfig, simulators, ParallelAdapter,
lazy diagnostics/output symbols).
Top-level package for gwmock-noise.
ARMANoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate stationary detector noise with a fitted ARMA / state-space model.
denominator_coefficients
property
¶
Return a copy of the ARMA denominator coefficients a_1 .. a_p.
metadata
property
¶
Return metadata describing the fitted ARMA model.
model_psd_curve
property
¶
Return the fitted model PSD on the band-masked design grid.
numerator_coefficients
property
¶
Return a copy of the ARMA numerator coefficients b_0 .. b_q.
resume_metadata_nbytes
property
¶
Serialized size of the resume metadata in bytes.
state_nbytes
property
¶
Serialized size of the continuation state in bytes.
target_psd_curve
property
¶
Return the band-masked target PSD on the design grid.
__init__(*, psd_file=None, target_psd=None, target_frequencies=None, ar_order=DEFAULT_AR_ORDER, ma_order=DEFAULT_MA_ORDER, line_frequencies=None, line_widths=None, detect_lines=False, max_lines=DEFAULT_MAX_LINES, line_prominence=DEFAULT_LINE_PROMINENCE, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization=DEFAULT_REGULARIZATION)
¶
Initialize the simulator and fit the ARMA model once.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_file
|
str | Path | None
|
Path or bundled preset name of a two-column PSD table. |
None
|
target_psd
|
ndarray | None
|
One-sided target PSD array, mutually exclusive with
|
None
|
target_frequencies
|
ndarray | None
|
Frequencies of |
None
|
ar_order
|
int
|
Autoregressive order |
DEFAULT_AR_ORDER
|
ma_order
|
int
|
Moving-average order |
DEFAULT_MA_ORDER
|
line_frequencies
|
list[float] | None
|
Explicit line frequencies to place poles at. |
None
|
line_widths
|
list[float] | None
|
Widths matching |
None
|
detect_lines
|
bool
|
Whether to detect prominent lines from the target. |
False
|
max_lines
|
int
|
Maximum number of detected lines. |
DEFAULT_MAX_LINES
|
line_prominence
|
float
|
Minimum peak-to-median ratio for detection. |
DEFAULT_LINE_PROMINENCE
|
detectors
|
list[str] | None
|
Detector names to generate. |
None
|
sampling_frequency
|
float
|
Sampling frequency in hertz. |
4096.0
|
duration
|
float
|
Nominal generation duration in seconds. |
4.0
|
seed
|
int | None
|
Base seed for the per-detector generators. |
None
|
low_frequency_cutoff
|
float
|
Lower edge of the target band in hertz. |
2.0
|
high_frequency_cutoff
|
float | None
|
Upper edge of the target band in hertz. |
None
|
block_size
|
int
|
Number of samples generated per filter call. |
DEFAULT_BLOCK_SIZE
|
regularization
|
float
|
Relative ridge added to the zero-lag autocovariance. |
DEFAULT_REGULARIZATION
|
continuation_state()
¶
Return the bounded state needed to keep a running stream going.
export_state()
¶
Return a picklable snapshot sufficient to resume the stream.
from_component(component, config)
classmethod
¶
Construct an ARMA simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector ARMA noise with continuity across calls.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield ARMA chunks lazily while preserving filter state.
import_state(state)
¶
Restore a snapshot produced by :meth:export_state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same configuration. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator's settings. |
reset()
¶
Clear filter state and RNGs.
resume_metadata()
¶
Return the metadata a stopped stream must persist to resume.
ARNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate stateful detector noise from an AR model fit to a target PSD.
autoregressive_coefficients
property
¶
Return a copy of the fitted a_1 .. a_p coefficients.
metadata
property
¶
Return metadata describing the fitted AR model.
model_psd_curve
property
¶
Return the fitted model PSD on the band-masked design grid.
resume_metadata_nbytes
property
¶
Serialized size of the resume metadata in bytes.
state_nbytes
property
¶
Serialized size of the continuation state in bytes.
target_psd_curve
property
¶
Return the band-masked target PSD on the design grid.
__init__(*, psd_file, order=DEFAULT_AR_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization=DEFAULT_REGULARIZATION)
¶
Initialize the simulator and fit the AR model once.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_file
|
str | Path
|
Path or bundled preset name of a two-column PSD table. |
required |
order
|
int
|
Autoregressive order |
DEFAULT_AR_ORDER
|
detectors
|
list[str] | None
|
Detector names to generate. |
None
|
sampling_frequency
|
float
|
Sampling frequency in hertz. |
4096.0
|
duration
|
float
|
Nominal generation duration in seconds. |
4.0
|
seed
|
int | None
|
Base seed for the per-detector generators. |
None
|
low_frequency_cutoff
|
float
|
Lower edge of the target band in hertz. |
2.0
|
high_frequency_cutoff
|
float | None
|
Upper edge of the target band in hertz. |
None
|
block_size
|
int
|
Number of samples generated per recursion call. |
DEFAULT_BLOCK_SIZE
|
regularization
|
float
|
Relative ridge added to the zero-lag autocovariance
before the recursion; |
DEFAULT_REGULARIZATION
|
continuation_state()
¶
Return the bounded state needed to keep a running stream going.
The state is the recursion memory of every detector; its size is proportional to the order and independent of the generated span.
export_state()
¶
Return a picklable snapshot sufficient to resume the stream.
from_component(component, config)
classmethod
¶
Construct an AR-noise simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector AR noise with continuity across calls.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield AR-noise chunks lazily while preserving recursion state.
import_state(state)
¶
Restore a snapshot produced by :meth:export_state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same configuration. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator's settings. |
reset()
¶
Clear detector state and RNGs.
resume_metadata()
¶
Return the metadata a stopped stream must persist to resume.
AddLines
¶
Wrap a base simulator and add spectral lines to its output.
metadata
property
¶
Return additive-wrapper metadata.
__init__(base, lines)
¶
Initialize the additive wrapper.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate base noise and add the configured spectral lines.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield additive spectral-line chunks lazily.
reset()
¶
Reset the additive line state and any resettable base state.
BaseNoiseSimulator
¶
Bases: ABC
Abstract base class for noise simulators.
This interface is the stable API through which the upstream gwmock package
interacts with gwmock_noise. Implementations must override :meth:run.
run(config)
abstractmethod
¶
Run the noise simulation with the given configuration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
NoiseConfig
|
Validated noise simulation configuration. |
required |
Returns:
| Type | Description |
|---|---|
SimulationResult
|
Result containing paths to generated outputs and the config used. |
BlipGlitch
dataclass
¶
Bases: GlitchModel
Gaussian-windowed broadband burst.
__post_init__()
¶
Validate blip-specific parameters.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate a Gaussian-windowed white-noise burst.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
ChannelMetadata
dataclass
¶
Typed, machine-checkable description of one channel.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Channel name, unique within one joint realization. |
kind |
ChannelKind
|
Whether the channel is a gravitational-wave strain channel or an auxiliary/witness channel. |
unit |
str
|
Physical unit of the channel's samples (e.g. |
domain |
ChannelDomain
|
Whether the channel's array is a time series or a frequency-domain series. |
sampling_frequency |
float
|
Sampling frequency in hertz. |
dtype |
str
|
The numpy dtype name (e.g. |
ColoredNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate colored detector noise from an input PSD.
metadata
property
¶
Return simulator metadata.
previous_strain
property
¶
Expose continuity buffers for protocol-compatible state inspection.
__init__(*, psd_file=None, psd_schedule=None, psd_array=None, frequencies=None, delta_frequency=None, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, window_duration=DEFAULT_WINDOW_DURATION)
¶
Initialize the simulator.
export_state()
¶
Return a picklable snapshot of the running crossfade stream.
The snapshot persists the bit-generator state and the chunk counter of the stitcher plus the bookkeeping needed to interpolate the PSD schedule on resume. It contains no cached strain: the previous window is regenerated from the bit-generator state.
from_component(component, config)
classmethod
¶
Construct a colored-noise simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector colored noise.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield colored-noise chunks lazily while preserving simulator state.
import_state(state)
¶
Resume the crossfade stream from an :meth:export_state snapshot.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same settings. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator. |
reset()
¶
Clear continuity and RNG state.
CompositeNoiseSimulator
¶
Additively combine multiple component simulators.
metadata
property
¶
Return metadata describing the composed run.
__init__(components, *, detectors, duration, sampling_frequency, seed)
¶
Initialize the composed simulator.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate all component realizations and add them together.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield additive component chunks lazily.
ConfigurableNoiseSimulator
¶
CorrelatedARNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate correlated detector noise with a truncated VMA representation.
metadata
property
¶
Return metadata describing the fitted VMA model.
__init__(*, psd_files, csd_files=None, order=DEFAULT_VMA_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization_epsilon=DEFAULT_REGULARIZATION_EPSILON)
¶
Initialize the simulator and fit truncated VMA taps once.
from_component(component, config)
classmethod
¶
Construct a correlated-AR simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate correlated per-detector noise with stateful continuity.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield correlated AR chunks lazily while preserving innovation state.
reset()
¶
Clear innovation state and RNG.
CorrelatedNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate correlated detector noise from PSD and CSD inputs.
metadata
property
¶
Return simulator metadata.
previous_strain
property
¶
Expose continuity buffers for protocol-compatible state inspection.
__init__(*, psd_files, csd_files=None, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, regularization_epsilon=1e-12, window_duration=DEFAULT_WINDOW_DURATION)
¶
Initialize the simulator.
from_component(component, config)
classmethod
¶
Construct a correlated-noise simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate correlated per-detector noise.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield correlated-noise chunks lazily while preserving simulator state.
reset()
¶
Clear continuity and RNG state.
DeepExtractorGlitch
dataclass
¶
Bases: GlitchModel
Real O3 glitch reconstructions colored against a target PSD.
Waveforms come from the DeepExtractor glitch-reconstruction dataset:
whitened, amplitude-normalized 2-second time series sampled at 4096 Hz,
covering seven Gravity Spy classes. Each drawn waveform is resampled to
the simulation rate, colored against the configured PSD, and rescaled so
its optimal SNR sqrt(4 df sum(|h(f)|^2 / S(f))) against that PSD
matches the configured target. The dataset (~2.3 GB) is downloaded lazily
on first use and cached by huggingface_hub; later runs reuse the cached
files. On each run the Hub is contacted first to validate the cached
files' ETags and download anything missing. If the Hub is unreachable, a
warning notes that the ETag check was skipped and the cached files are used
instead; if they are not cached, the first generate_waveform raises
huggingface_hub's LocalEntryNotFoundError (a FileNotFoundError
subclass). Set local_files_only=True to skip the network unconditionally
and read straight from the cache.
revision pins the download to a specific dataset version — a git branch,
tag, or commit SHA forwarded to hf_hub_download. Leave it None to
track the repository default. Either way, the first download resolves to a
concrete commit SHA (read back from the Hub cache path), which is what
serialize records. Replaying a run through its metadata therefore fetches
the exact commit that produced it, keeping glitch generation bit-reproducible
for a fixed (version, config, seed) even as the upstream dataset moves.
rate accepts either a single number — the total Poisson rate shared by
all configured classes, drawn uniformly — or a mapping from class name to
per-class Poisson rate, in which case the total rate is their sum and each
event's class is drawn proportionally to its rate.
snr accepts a number, a mapping from class name to number, a distribution,
or a mapping from class name to distribution — freely mixed, so one class can
be sampled while another stays fixed. A number calibrates every event of that
class to exactly that loudness; a distribution draws a target per event from
the model's own random stream, which is what a measured glitch population
needs, since each Gravity Spy class is heavy-tailed and the tail indices
differ by close to an order of magnitude between classes. See
:mod:gwmock_noise.glitches.snr for the supported shapes.
What the target means is the caller's to decide. A tail index measured
from a trigger generator's SNR — Omicron's, say — is defined against that
pipeline's PSD over the band it searched, while this model calibrates against
psd_file from low_frequency_cutoff upward. Configuring one from the
other equates two SNRs defined against different noise curves over different
bands. The sampler reproduces whatever distribution it is given and cannot
tell whether that identification is the one you meant.
Resampling uses linear interpolation without an anti-aliasing filter, so sampling frequencies below 4096 Hz alias high-frequency content; the SNR calibration itself is unaffected because it is computed after resampling.
__post_init__()
¶
Validate the configuration and preload the PSD table.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate one colored, SNR-calibrated DeepExtractor glitch.
resolve()
¶
Pin the dataset to a concrete commit SHA, fetching only the samples file.
Reads the commit SHA huggingface_hub stamps into the cache path so the
resolved revision is available for reproducibility metadata even when no
waveform has been drawn yet — e.g. a batch that fires zero glitch events,
where generate_waveform (and therefore _get_dataset) never runs.
The samples file is downloaded once and reused by the later full load, so
this adds no extra download on the generation path. Returns the resolved
SHA, or None when the cache layout does not expose one (e.g. a test
stub serving files from a flat directory).
serialize()
¶
Return metadata-friendly model parameters.
DefaultNoiseSimulator
¶
Bases: BaseNoiseSimulator
Default noise simulator implementation.
metadata
property
¶
Return metadata describing the current simulator state.
__init__(*, duration=4.0, sampling_frequency=4096.0, detectors=None, seed=None)
¶
Initialize the simulator with protocol-compatible state.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Return Gaussian white-noise strain arrays.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield white-noise strain chunks lazily.
run(config)
¶
Run the noise simulation with the given configuration.
DiagnosticResult
dataclass
¶
Outcome of a statistical diagnostic check.
EmpiricalSNRDistribution
dataclass
¶
Bases: SNRDistribution
Target SNR drawn with replacement from a table of observed SNRs.
For a caller holding the measurements rather than a fit to them: the drawn
population is exactly the supplied one, tail included, with no assumption
about its shape. Supply the values inline as samples or point file
at a table on disk -- exactly one of the two.
Attributes:
| Name | Type | Description |
|---|---|---|
samples |
Sequence[float] | None
|
Observed SNRs given inline, or |
file |
str | Path | None
|
Path to a table of observed SNRs, or |
distribution |
str
|
Discriminator naming this shape in a configuration. |
__post_init__()
¶
Validate the configuration and load the SNR table.
sample(rng)
¶
Draw one target SNR uniformly from the table, with replacement.
serialize()
¶
Return the mapping that reconstructs this distribution.
Echoes whichever of the two inputs was configured, so a file-backed table replays as a reference to the file rather than as a copy of its contents inlined into the metadata.
survival(threshold)
¶
Return the fraction of the supplied table at or above threshold.
ExpectedCount
dataclass
¶
The expected injected count for one class in one interferometer, factorized.
The product is :attr:expected, and every factor is carried beside it. A campaign that
finds the total surprising can then read off which factor is surprising -- a rate
transplanted from another run, a livetime that is calendar time rather than observing
time, a participation probability nobody meant to set -- instead of re-deriving the
whole chain.
Attributes:
| Name | Type | Description |
|---|---|---|
glitch_class |
str
|
Name of the class within its population. |
detector |
str
|
The interferometer this row is for. |
rate_hz |
float
|
The class's Poisson rate. For a class whose interferometers glitch independently this is the rate each of them sees; for a network-coherent class it is the rate of the shared network process, which is not the same number unless every interferometer takes every event. |
livetime_seconds |
float
|
The analysed time this count is over. |
participation |
float
|
Probability that this interferometer receives a given event of the class. One for a class with no network coherence. |
selection |
float
|
Fraction of the class's events that survive the SNR cut the count was asked for. One when no cut was given. |
expected |
float
|
|
network_events
property
¶
Expected number of network events of this class over the same livetime.
The same arithmetic with the participation factor divided out: the events the class's process produced, rather than the subset this interferometer received. For a class with no network coherence the two coincide, because every event it draws for an interferometer is that interferometer's.
FitError
¶
Bases: ValueError
Raised when a model fit is numerically invalid or unstable.
GengliBlipGlitch
dataclass
¶
Bases: GlitchModel
File-backed gengli blip generator colored against a target PSD.
__post_init__()
¶
Validate configured files and preload the population/PSD tables.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
from_population_file(population_file, **kwargs)
classmethod
¶
Construct a gengli glitch model from a population file path.
generate_waveform(sampling_frequency, rng=None)
¶
Generate one colored gengli blip waveform.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
GlitchClass
dataclass
¶
One named class within a registered population.
The prose fields are part of the registered definition and therefore part of the digest, deliberately. A population whose numbers are unchanged but whose meaning has been rewritten is a different population, and a digest that did not move would say otherwise.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Stable identifier of the class within its population. |
morphology |
str
|
Short slug naming the waveform family, e.g.
|
description |
str
|
One sentence saying what the class represents and what it is drawn from. |
model |
GlitchModel
|
The glitch model that generates it, already normalized. |
__post_init__()
¶
Validate the class definition.
Raises:
| Type | Description |
|---|---|
ValueError
|
If the name, morphology or description is blank. |
deserialize(payload)
classmethod
¶
Rebuild a class from :meth:serialize output.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
payload
|
dict[str, Any]
|
The serialized class. |
required |
Returns:
| Type | Description |
|---|---|
GlitchClass
|
The reconstructed class. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If a required key is missing. |
serialize()
¶
Return the mapping that reconstructs this class.
GlitchDraw
¶
Bases: NamedTuple
One drawn glitch waveform and the parameters that produced it.
What :meth:GlitchModel.draw returns so an injector can record what it
injected, not merely that it injected something. Every field but the
waveform is optional because not every model has the quantity: a parametric
blip with no PSD has no SNR to report, and only the dataset-backed models
draw a Gravity Spy class.
Attributes:
| Name | Type | Description |
|---|---|---|
waveform |
ndarray
|
The glitch strain, sampled at the requested rate. Its first sample lands at the injection time -- the waveform starts there, it does not peak there. |
amplitude |
float | None
|
The amplitude multiplier drawn from the model's amplitude
distribution, or |
glitch_class |
str | None
|
The drawn morphological class (e.g. a Gravity Spy class
name), or |
target_snr |
float | None
|
The optimal SNR the draw was calibrated to before the
amplitude multiplier, or |
realized_snr |
float | None
|
The optimal SNR of the whole of How it relates to |
GlitchModel
dataclass
¶
Base dataclass for transient glitch generators.
detectors scopes the model to a subset of the run's interferometers. It
accepts one name or a list of them, and defaults to None, which is every
interferometer -- what a model without a selector has always meant. Scoping is
what lets one configuration carry a 10 km model for one set of interferometers
and a 15 km model for another, and per-interferometer rates for a network whose
instruments are not equally glitchy (in O3, Fast_Scattering fired about 29 times
more often in L1 than in H1).
A selector is only half of it. :func:validate_glitch_detector_coverage refuses
a run whose models do not between them account for every interferometer, or
whose models color one interferometer against two different PSDs -- see its
docstring for why proceeding was the defect.
__post_init__()
¶
Validate common glitch parameters.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate a single glitch waveform.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
GlitchNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate transient glitches as an additive standalone component.
from_component(component, config)
classmethod
¶
Construct a glitch-only component from one component definition.
GlitchPopulation
dataclass
¶
A named, versioned, digestible set of glitch classes.
The digest is only as portable as what goes into it. A model's psd_file is
serialized as written, so a population built against /scratch/me/et.txt carries
that path into its digest and no other machine can reproduce it. Name a bundled preset,
or a path relative to the campaign's own tree, and the digest travels.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Stable registered identifier, e.g. |
description |
str
|
One paragraph saying what the population is for and what it is anchored to. |
classes |
tuple[GlitchClass, ...]
|
The classes, in registration order. |
references |
tuple[str, ...]
|
Published sources the registered quantities are anchored to, one citation per entry. |
unanchored |
tuple[str, ...]
|
Quantities that could not be anchored to a published measurement, one per entry, each saying what it is and why. Carried in the population rather than only in its documentation, so a consumer reading a pinned population reads the caveats with it. An empty tuple is a claim that everything is anchored, so it is a claim to be able to defend. |
schema_version |
str
|
Serialization schema, part of the digest. |
__post_init__()
¶
Validate the population definition.
Raises:
| Type | Description |
|---|---|
ValueError
|
If the name or description is blank, the population has no classes, or two classes share a name. |
build_injector(base, *, gps_start=0.0)
¶
Wrap a base simulator so it also injects this population.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
base
|
NoiseSimulator
|
The simulator producing the noise the glitches are added to. |
required |
gps_start
|
float
|
GPS time of the first sample, stamped onto the truth catalogue. |
0.0
|
Returns:
| Type | Description |
|---|---|
InjectGlitches
|
The wrapped simulator. |
canonical_json()
¶
Return the exact bytes the digest is taken over.
Sorted keys, no insignificant whitespace, ASCII only and no NaN: a mapping that
two processes can be relied on to render identically. Exposed rather than hidden so
a campaign that disagrees about a digest can diff what was hashed instead of
guessing.
Returns:
| Type | Description |
|---|---|
str
|
The canonical serialization. |
deserialize(payload)
classmethod
¶
Rebuild a population from :meth:serialize output.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
payload
|
dict[str, Any]
|
The serialized population. |
required |
Returns:
| Type | Description |
|---|---|
GlitchPopulation
|
The reconstructed population, which serializes and digests identically. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If a required key is missing. |
ValueError
|
If the payload declares a schema version this package cannot read. |
digest()
¶
Return the population's sha256:... digest.
What a campaign manifest pins. Every registered quantity, every prose field and the schema version go into it, so a population that has been edited cannot present itself as the one a finished run used.
Returns:
| Type | Description |
|---|---|
str
|
The digest, algorithm-prefixed. |
expected_counts(*, livetime_seconds, detectors, snr_threshold=None)
¶
Decompose the expected injected count per class and interferometer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
livetime_seconds
|
float
|
Analysed time the count is over. This is observing time, not calendar time: the distinction is the commonest factor-of-1.3 error in a count derived from a published rate. |
required |
detectors
|
Sequence[str]
|
The run's interferometers. Checked against the population's own scoping, so a decomposition cannot be computed for a network the population does not account for. |
required |
snr_threshold
|
float | None
|
Optional recording or detection cut. When given, the selection factor is the fraction of each class's target-SNR distribution at or above it, read off the distribution in closed form. |
None
|
Returns:
| Type | Description |
|---|---|
list[ExpectedCount]
|
One row per (class, interferometer) pair, in registration order. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the livetime is not positive, no interferometers were given, or the population does not account for every interferometer. |
model_for(glitch_class)
¶
Return one class's glitch model by name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
glitch_class
|
str
|
The class name. |
required |
Returns:
| Type | Description |
|---|---|
GlitchModel
|
The model that generates it. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If the population has no such class. |
models()
¶
Return the population's glitch models in registration order.
realize(*, detectors, duration, sampling_frequency, seed, gps_start=0.0)
¶
Generate this population's glitch strain on its own, with no noise under it.
This is the object paired arms share. A comparison whose arms differ in their
signal content, their geometry, their detection threshold or their ranking
statistic must not differ in their glitches, and the way to guarantee that is for
every arm to add the same array rather than to re-run a generator and trust it to
agree. The returned stamp carries the population digest and the seed, so an arm can
record which realization it added -- and the digest is the field to pin, for the
reason :class:GlitchRealization gives.
The realization is reproducible for a fixed (package version, population, seed) and does not depend on the base noise, on the other interferometers in the run, or on whether it was generated in one call or streamed -- see the module's tests, which assert each of those separately.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
Sequence[str]
|
The interferometers to realize. |
required |
duration
|
float
|
Length of the realization in seconds. |
required |
sampling_frequency
|
float
|
Sampling frequency in Hz. Note that a class with a
|
required |
seed
|
int
|
The realization's seed. |
required |
gps_start
|
float
|
GPS time of the first sample. |
0.0
|
Returns:
| Type | Description |
|---|---|
GlitchRealization
|
The glitch strain, its truth catalogue and the stamp identifying it. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the population does not account for every interferometer. |
serialize()
¶
Return the mapping that reconstructs this population.
GlitchRealization
¶
Bases: NamedTuple
One seeded glitch realization and the stamp that identifies it.
Attributes:
| Name | Type | Description |
|---|---|---|
strain |
dict[str, ndarray]
|
The glitch contribution per interferometer, with nothing else in it. This is what paired arms add to their own base noise, and adding the same array to two arms is what makes their glitch content bit-identical. |
events |
list[dict[str, Any]]
|
The per-event truth catalogue, in time order. See
:data: |
stamp |
dict[str, Any]
|
What identifies this realization: the population's name and digest, the
package version, and the generation parameters. Enough to reproduce The pinned value is |
GwoscFilterConfig
¶
Bases: BaseModel
Configuration for filtering GWOSC strain segments.
Attributes:
| Name | Type | Description |
|---|---|---|
filter_types |
list[FilterType]
|
Which categories of segments to exclude. |
far_threshold |
float
|
Maximum false-alarm rate (events/year) for GW event filtering. Events with FAR above this threshold are excluded from filtering (treated as noise). Default 1.0 (one per year). |
event_padding |
float
|
Padding in seconds to apply around each GW event GPS time when building vetosegments. |
dq_flags |
list[str]
|
DQ flag basenames (without detector prefix) to query for
data-quality vetosegments. E.g. |
exclude_hardware_injections |
bool
|
Whether to also exclude segments containing hardware injections. |
GwoscNoiseConfig
¶
Bases: BaseModel
Configuration for fetching real detector noise from GWOSC.
Attributes:
| Name | Type | Description |
|---|---|---|
detectors |
list[str]
|
List of detector prefixes (e.g. |
gps_start |
float
|
GPS start time of the requested data interval. |
gps_end |
float
|
GPS end time of the requested data interval. |
sample_rate |
float
|
Desired sampling rate in Hz. GWOSC typically provides 4096 Hz, with some event datasets at 16384 Hz. |
filters |
GwoscFilterConfig
|
Filter configuration for excluding segments. |
host |
str
|
GWOSC host URL. |
cache_dir |
Path | None
|
Optional directory for file-level caching. When set,
downloaded HDF5 frame files are saved locally and reused on
subsequent requests for the same GPS interval. When |
GwoscNoiseFetcher
¶
Fetch real detector noise data from GWOSC with optional filtering.
Downloads strain data from GWOSC and applies user-configured filters to exclude segments containing GW signals and data-quality issues.
When cache_dir is configured, HDF5 files are saved locally and
reused on subsequent requests — avoiding repeated downloads for the
same GPS interval.
Attributes:
| Name | Type | Description |
|---|---|---|
config |
The GWOSC noise fetching configuration. |
clean_segments
property
¶
__init__(config)
¶
Initialize the fetcher.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
GwoscNoiseConfig
|
Configuration specifying detectors, GPS range, sample rate, filtering options, and optional cache directory. |
required |
check_availability(detectors=None)
¶
Probe GWOSC for per-detector strain-data availability.
Unlike :attr:clean_segments, which only reflects vetoes (GW
events and data-quality flags), this checks whether the strain
data itself has been published and is downloadable for the full
configured GPS interval. A detector can have a fully "clean"
interval yet have no published strain — e.g. the data has not
been released yet for the observing run — in which case a fetch
would fail. Use this for a pre-flight check before fetching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str] | None
|
Detectors to probe. Defaults to all configured
detectors when |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, bool]
|
A dictionary mapping each requested detector to |
dict[str, bool]
|
GWOSC has data URLs covering the interval, |
fetch_clean(detectors=None)
¶
Fetch clean noise segments.
Clean segments are computed by excluding GW events and data-quality issues according to the filter configuration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str] | None
|
Detectors to fetch. Defaults to all configured
detectors when |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, list[TimeSeries]]
|
A dictionary mapping each requested detector to a list of |
dict[str, list[TimeSeries]]
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If no data is available for any requested detector or no clean segments are found. |
fetch_raw(detectors=None)
¶
Fetch raw strain data without filtering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str] | None
|
Detectors to fetch. Defaults to all configured
detectors when |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, TimeSeries]
|
A dictionary mapping each requested detector to a full-interval |
dict[str, TimeSeries]
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If no data is available for any requested detector. |
GwoscNoiseSimulator
¶
Real-noise simulator that fetches strain data from GWOSC.
Implements the NoiseSimulator protocol so it can be used
interchangeably with synthetic simulators like
ColoredNoiseSimulator and CorrelatedNoiseSimulator.
Fetches real detector strain from the Gravitational-Wave Open Science Centre and applies user-configured filters to exclude segments containing GW signals or data-quality issues.
When cache_dir is set in the config, downloaded HDF5 files
are saved locally and reused — avoiding repeated downloads for
the same GPS interval.
Attributes:
| Name | Type | Description |
|---|---|---|
duration |
float
|
Duration of the configured GPS interval (seconds). |
sampling_frequency |
float
|
Sampling frequency in Hz. |
detectors |
list[str]
|
List of detector prefixes. |
seed |
None
|
Always |
config |
The underlying GWOSC configuration. |
detectors
property
¶
Return the list of detector prefixes.
duration
property
¶
Return the total duration of the configured GPS interval.
metadata
property
¶
sampling_frequency
property
¶
Return the sampling frequency in Hz.
seed
property
¶
Return None — real noise has no controllable random seed.
__init__(config)
¶
Initialize the real-noise simulator.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
GwoscNoiseConfig
|
Configuration specifying GPS range, detectors, sample rate, filtering options, and optional cache directory. |
required |
check_availability(detectors=None)
¶
Probe GWOSC for per-detector strain-data availability.
Convenience passthrough to :meth:GwoscNoiseFetcher.check_availability.
Use this as a pre-flight check: a detector may have a fully clean
interval (no vetoes) yet have no published strain data, in which
case :meth:generate would fail with a clear error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str] | None
|
Detectors to probe. Defaults to all configured
detectors when |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, bool]
|
A dictionary mapping each requested detector to |
dict[str, bool]
|
GWOSC has data covering the interval, |
generate(duration, sampling_frequency, detectors, seed=None)
¶
Fetch clean noise and return the longest contiguous segment per detector.
Real GWOSC noise is fragmented by the excision of GW events and
data-quality vetoes, so the clean data generally consists of
several non-contiguous spans. To honour the NoiseSimulator
contract — a single, gap-free array per detector — this returns
the longest contiguous clean segment for each detector.
Concatenating the fragments instead (the historical behaviour)
splices non-adjacent samples together, creating discontinuities
that show up as sharp transients after whitening. Use
:meth:generate_segments to obtain every clean segment
separately when the full clean data is required.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
duration
|
float
|
Requested duration (ignored; the longest available contiguous clean segment determines the output length). |
required |
sampling_frequency
|
float
|
Requested sampling frequency (must match
the configured |
required |
detectors
|
list[str]
|
Requested detector list (must be a subset of the configured detectors). |
required |
seed
|
int | None
|
Ignored for real noise. |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, ndarray]
|
A dictionary mapping each detector to a 1-D numpy array |
dict[str, ndarray]
|
holding its longest contiguous clean segment. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
generate_segments(duration, sampling_frequency, detectors, seed=None)
¶
Fetch clean noise and return per-detector lists of contiguous segments.
Unlike :meth:generate, this preserves the gap structure of the
real data: each clean segment — the spans left after excluding GW
events and data-quality vetoes — is returned as its own array, in
GPS-time order. Adjacent segments are not contiguous in time,
so callers must treat each segment independently; concatenating
them would splice non-adjacent samples together and introduce
discontinuities at the joins (which manifest as sharp transients
after whitening).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
duration
|
float
|
Requested duration (ignored; the GPS interval from the config determines the available data). |
required |
sampling_frequency
|
float
|
Requested sampling frequency (must match
the configured |
required |
detectors
|
list[str]
|
Requested detector list (must be a subset of the configured detectors). |
required |
seed
|
int | None
|
Ignored for real noise. |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, list[ndarray]]
|
A dictionary mapping each detector to a list of 1-D numpy |
dict[str, list[ndarray]]
|
arrays, one per clean segment, ordered by GPS time. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield clean-noise chunks lazily.
Fetches the longest contiguous clean segment once (see
:meth:generate) and yields it in chunks of chunk_duration
seconds. Chunking the longest contiguous span keeps every chunk
free of the discontinuities that splicing fragments would
introduce.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
chunk_duration
|
float
|
Duration of each yielded chunk in seconds. |
required |
sampling_frequency
|
float
|
Requested sampling frequency (must
match the configured |
required |
detectors
|
list[str]
|
Requested detector list (must be a subset of the configured detectors). |
required |
seed
|
int | None
|
Ignored for real noise. |
None
|
Yields:
| Type | Description |
|---|---|
dict[str, ndarray]
|
Per-detector strain arrays for each chunk. |
GwoscSegmentFilter
¶
Compute clean analysis segments from GWOSC metadata.
Queries GWTC event catalogs and data-quality flags to build vetosegments (time windows to exclude) and returns the remaining clean intervals suitable for noise analysis.
Attributes:
| Name | Type | Description |
|---|---|---|
config |
The filter configuration. |
__init__(config)
¶
Initialize the segment filter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
GwoscFilterConfig
|
Filter configuration specifying which segment categories to exclude and their parameters. |
required |
compute_clean_segments(gps_start, gps_end, detectors)
¶
Compute clean (analysis-ready) segments for each detector.
Clean segments are the requested GPS range minus the union of: - GW event vetosegments (if any GW filter is active) - DQ vetosegments (if the data-quality filter is active) - Hardware injection segments (if enabled)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gps_start
|
float
|
GPS start of the requested interval. |
required |
gps_end
|
float
|
GPS end of the requested interval. |
required |
detectors
|
list[str]
|
List of detector prefixes. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, SegmentList]
|
A dictionary mapping each detector to a list of clean |
dict[str, SegmentList]
|
|
get_dq_vetosegments(gps_start, gps_end, detector)
¶
Query data-quality flags for detector and build vetosegments.
Uses gwosc.timeline.get_segments() to fetch pre-computed
DQ veto segments for each flag in dq_flags.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gps_start
|
float
|
GPS start of the query interval. |
required |
gps_end
|
float
|
GPS end of the query interval. |
required |
detector
|
str
|
Detector prefix (e.g. |
required |
Returns:
| Type | Description |
|---|---|
SegmentList
|
A list of |
SegmentList
|
time windows with data-quality issues. |
get_gw_vetosegments(gps_start, gps_end)
¶
Query GWTC events in the GPS range and build vetosegments.
For HIGH_CONFIDENCE_GW, only events with FAR <= far_threshold
are included. For ALL_GW_SIGNALS, all events in the range are
included.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gps_start
|
float
|
GPS start of the query interval. |
required |
gps_end
|
float
|
GPS end of the query interval. |
required |
Returns:
| Type | Description |
|---|---|
SegmentList
|
A list of |
SegmentList
|
on a GW event GPS time with event_padding on both sides. |
get_hardware_injection_segments(gps_start, gps_end)
¶
Query hardware injection events and build vetosegments.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gps_start
|
float
|
GPS start of the query interval. |
required |
gps_end
|
float
|
GPS end of the query interval. |
required |
Returns:
| Type | Description |
|---|---|
SegmentList
|
A list of |
SegmentList
|
hardware injection times. |
InjectGlitches
¶
Wrap a base simulator and inject transient glitches additively.
A glitch model with no network coherence runs an independent Poisson process
per detector, so every detector receives its own event times and waveform
realizations, and GlitchModel.rate is the event rate seen by each individual
detector. Every (model, detector) pair draws from its own random generator,
derived from the top-level seed and the detector name via an independent
SeedSequence stream.
A model carrying a :class:~gwmock_noise.glitches.models.NetworkCoherence runs
one Poisson process for the whole network and draws one waveform per
event, offering it to every detector it applies to. GlitchModel.rate is then
the network event rate, not the per-detector one, and each detector sees it
times the coherence's participation probability. Arrival times and waveforms come
from a stream keyed by the model alone; each detector's participation and
amplitude ratio come from a stream keyed by its own name and the event's ordinal.
Either way, a detector's realization is reproducible and independent of which other detectors are present or the order in which they are requested -- which is what lets two runs over different networks be compared on the channels they share.
A glitch whose waveform extends past the end of a chunk has its remainder carried into the next chunk and replayed in event order, so the injected glitch series is identical, sample for sample, to a single generate call of the same total duration (the per-sample addition order is preserved exactly, not merely up to rounding). A tail is dropped where the data it would continue into does not exist: past the last generated chunk, and past a segment whose successor the caller places at a discontinuous epoch.
Every event that lands in the strain is recorded as it fires, not merely
tallied: see :attr:glitch_events for the cumulative truth catalogue and
:attr:segment_glitch_events for the events the most recent generate call
injected. A row is written where a glitch starts, so a waveform straddling
a chunk boundary appears once, in the chunk that holds its first sample, at
its true time -- the carried-over tail adds samples to the next chunk but no
second row. GLITCH_CATALOGUE_COLUMNS documents the columns and
GLITCH_CATALOGUE_TIME_CONVENTION which instant each time column names.
gps_start is the GPS time of the first sample of the next segment, and it
belongs to the caller. It advances by each generated segment's duration, so a
stream's catalogue reads in real GPS time on its own; a caller that knows
better -- a writer placing non-contiguous segments, or one driving the stream
batch by batch -- assigns it before each generate call and that assignment is
honoured, including on a call that also carries a seed. Only reset, the
explicit start-over, rewinds it to the epoch the wrapper was constructed with.
glitch_events
property
¶
Return every glitch injected since the last (re)seed, in time order.
Copies, so a consumer cannot edit the run's own record of itself.
metadata
property
¶
Return additive-wrapper metadata.
segment_glitch_events
property
¶
Return the glitches the most recent generate call injected.
What a per-batch consumer wants: the rows belonging to the chunk it is about to write, rather than the whole stream's catalogue re-filtered by time. Empty before the first generate call, and after a call that fired no events.
__init__(base, glitch_models, gps_start=0.0)
¶
Initialize the additive glitch wrapper.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate base noise and inject the configured glitches.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield glitch-injected chunks lazily.
reset()
¶
Reset the additive wrapper and any resettable base state.
Rewinds the epoch to the one the wrapper was constructed with, which
_initialize_process does not: this is the caller saying "start over", where a
seed passed to generate is the caller saying "this segment, whose epoch I have
just told you, begins a new realization". The two need different answers, and
conflating them is what put every streamed segment on the wrong epoch.
JointCovariance
dataclass
¶
A one-sided cross-spectral density matrix over a set of channels.
The convention matches :func:numpy.fft.rfftfreq: frequencies is the
one-sided, non-negative frequency grid, and matrix[k] is the
Hermitian (n_channels, n_channels) one-sided cross-spectral density at
frequencies[k], in units of channel_unit**2 / Hz. The diagonal
entry matrix[k, i, i] is the one-sided PSD of channel
channel_order[i] at that frequency.
Attributes:
| Name | Type | Description |
|---|---|---|
frequencies |
ndarray
|
One-sided frequency grid in hertz, |
matrix |
ndarray
|
Cross-spectral density matrices, |
channel_order |
list[str]
|
Channel names, in the order |
JointDummyCorrelatedSimulator
¶
Public dummy backend producing simultaneous, correlated strain and witness channels.
Every channel is temporally white with unit per-sample variance and
coupling correlation with every other channel; there is no spectral
shaping and no physical model. The one-sided cross-spectral density is
therefore flat and known in closed form: for per-sample covariance
Sigma at sampling frequency fs, the one-sided PSD/CSD is the
frequency-independent matrix 2 * Sigma / fs (the same
covariance-to-PSD scale convention already used by
:class:~gwmock_noise.simulators.multichannel.MultichannelNoiseSimulator,
inverted).
channel_metadata
property
¶
Return typed metadata for every strain and witness channel.
metadata
property
¶
Return simulator metadata.
__init__(*, detectors, witnesses, sampling_frequency=DEFAULT_SAMPLING_FREQUENCY, duration=DEFAULT_DURATION, seed=None, coupling=DEFAULT_COUPLING, strain_unit='strain', witness_unit='counts')
¶
Initialize the dummy backend.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str]
|
Strain channel names. Fixed for the lifetime of this instance; later calls may reorder but not change this set. |
required |
witnesses
|
list[str]
|
Witness channel names, disjoint from |
required |
sampling_frequency
|
float
|
Sampling frequency in hertz. |
DEFAULT_SAMPLING_FREQUENCY
|
duration
|
float
|
Nominal generation duration in seconds. |
DEFAULT_DURATION
|
seed
|
int | None
|
Seed for the shared innovation generator. |
None
|
coupling
|
float
|
Equicorrelation coefficient applied between every pair
of channels (strain-strain, witness-witness and
strain-witness alike). Must satisfy |
DEFAULT_COUPLING
|
strain_unit
|
str
|
Unit string recorded for strain channels. |
'strain'
|
witness_unit
|
str
|
Unit string recorded for witness channels. |
'counts'
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
covariance()
¶
Return this backend's exact, closed-form one-sided cross-spectral density.
Every channel is temporally white, so the returned matrix is
frequency-independent: 2 * Sigma / sampling_frequency at every
frequency, where Sigma is the per-sample covariance matrix drawn
from at construction. This is an exact analytic value, not a
statistical estimate off a realization.
generate_joint(duration, sampling_frequency, detectors, witnesses, seed=None)
¶
Generate one simultaneous, correlated strain+witness realization.
generate_joint_stream(chunk_duration, sampling_frequency, detectors, witnesses, seed=None)
¶
Yield simultaneous strain+witness chunks lazily, preserving RNG state.
JointRealization
dataclass
¶
One simultaneous draw of strain and witness channels.
Attributes:
| Name | Type | Description |
|---|---|---|
strain |
dict[str, ndarray]
|
Per-detector strain arrays, one entry per requested detector. |
witness |
dict[str, ndarray]
|
Per-channel witness arrays, one entry per requested witness channel. |
channel_metadata |
dict[str, ChannelMetadata]
|
Typed metadata for every channel in |
provenance |
dict[str, Any]
|
Implementation-defined provenance describing how this realization was produced (backend identity, package version, seed, RNG algorithm, and any upstream configuration actually used). |
JointStrainWitnessSimulator
¶
Bases: Protocol
Structural interface for simulators producing joint strain+witness output.
This is an optional extension of the legacy NoiseSimulator contract:
implementing it changes nothing about NoiseSimulator itself, and a
NoiseSimulator-only consumer is unaffected by its existence. A class
may satisfy both protocols at once.
metadata
property
¶
Return simulator metadata (same role as NoiseSimulator.metadata).
covariance()
¶
Return the joint one-sided cross-spectral density over all channels.
generate_joint(duration, sampling_frequency, detectors, witnesses, seed=None)
¶
Generate one simultaneous strain+witness realization.
generate_joint_stream(chunk_duration, sampling_frequency, detectors, witnesses, seed=None)
¶
Yield simultaneous strain+witness chunks lazily.
Continuity contract: consecutive chunks from this iterator must equal
the same realization a caller would obtain from one seeded
:meth:generate_joint call spanning the combined duration -- the
same contract NoiseSimulator.generate_stream makes for
strain-only output.
LogNormalAmplitudeDistribution
dataclass
¶
MultichannelNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate multichannel noise from a tabulated PSD/CSD matrix.
ar_coefficients
property
¶
Return a copy of the fitted coefficient matrices A_1 .. A_p.
design_frequencies
property
¶
Return a copy of the design-grid frequencies in hertz.
innovation_covariance
property
¶
Return a copy of the fitted innovation covariance.
metadata
property
¶
Return metadata describing the fitted matrix spectral factor.
model_spectral_matrices
property
¶
Return a copy of the fitted cross-spectral matrices on the design grid.
model_spectral_matrix_curve
property
¶
Return the band-masked fitted cross-spectral matrices on the design grid.
resume_metadata_nbytes
property
¶
Serialized size of the resume metadata in bytes.
state_nbytes
property
¶
Serialized size of the continuation state in bytes.
target_spectral_matrices
property
¶
Return a copy of the target cross-spectral matrices on the design grid.
The matrices are in the covariance scale the recursion uses
(one_sided_psd * sampling_frequency / 2) and carry the band mask and
edge taper the fit was built from.
target_spectral_matrix_curve
property
¶
Return the band-masked target cross-spectral matrices on the design grid.
__init__(*, psd_files=None, csd_files=None, target_matrices=None, target_frequencies=None, order=DEFAULT_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization_epsilon=DEFAULT_REGULARIZATION_EPSILON)
¶
Initialize the simulator and fit the matrix spectral factor once.
The cross-spectral matrix may be given either as files (psd_files
plus csd_files, absent off-diagonal pairs mean zero coherence) or as
an in-memory target_matrices array on target_frequencies, which
lets an analytic pair bypass the tabulated-curve interpolation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_files
|
dict[str, str | Path] | None
|
Per-detector PSD table paths, keyed by detector. |
None
|
csd_files
|
dict[str, str | Path] | dict[tuple[str, str], str | Path] | None
|
Cross-spectral table paths, keyed by detector pair or by
|
None
|
target_matrices
|
ndarray | None
|
One-sided cross-spectral matrices of shape
|
None
|
target_frequencies
|
ndarray | None
|
Frequencies of |
None
|
order
|
int
|
Order |
DEFAULT_ORDER
|
detectors
|
list[str] | None
|
Detector names to generate. |
None
|
sampling_frequency
|
float
|
Sampling frequency in hertz. |
4096.0
|
duration
|
float
|
Nominal generation duration in seconds. |
4.0
|
seed
|
int | None
|
Seed for the shared innovation generator. |
None
|
low_frequency_cutoff
|
float
|
Lower edge of the fit band in hertz. |
2.0
|
high_frequency_cutoff
|
float | None
|
Upper edge of the fit band in hertz. |
None
|
block_size
|
int
|
Number of samples generated per recursion call. |
DEFAULT_BLOCK_SIZE
|
regularization_epsilon
|
float
|
Relative ridge, as a fraction of the largest in-band diagonal entry, added as a zero-lag (white) floor to the autocovariance when the in-band target is not positive definite. Zero forbids the ridge and makes such a target a fit failure. |
DEFAULT_REGULARIZATION_EPSILON
|
continuation_state()
¶
Return the bounded state needed to keep a running stream going.
The state is the recursion history plus the pending output; its size is proportional to the order and independent of the generated span.
export_state()
¶
Return a picklable snapshot sufficient to resume the stream.
from_component(component, config)
classmethod
¶
Construct a multichannel simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector multichannel noise with continuity across calls.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield multichannel chunks lazily while preserving recursion history.
import_state(state)
¶
Restore a snapshot produced by :meth:export_state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same configuration. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator's settings. |
reset()
¶
Clear the recursion history, pending output and generator.
resume_metadata()
¶
Return the metadata a stopped stream must persist to resume.
NetworkCoherence
dataclass
¶
Share one glitch model's events across the interferometers it applies to.
Without it, every glitch model runs an independent Poisson process and an independent random stream per interferometer, so two interferometers never carry the same transient. That is the right default, and it is what the one measurement of the question says for widely separated sites: over O1 and O2 the LIGO blip population produced no coincidences inside the ±15 ms window an astrophysical signal can occupy, and the coincidences found in wider windows matched the accidental expectation.
It is not what a network of co-located interferometers is expected to do. Three interferometers sharing one site, one vacuum system and one seismic environment see a common environmental transient in all three at once, and a coincidence veto -- or a null stream -- assumes exactly the incoherence that such an event breaks. A model carrying this runs one Poisson process for the whole network, draws one waveform per event, and offers that waveform to each interferometer it applies to.
Participation is decided per interferometer, independently, and is never conditioned on how many others took the event. Conditioning -- "keep only events that land in at least two interferometers" -- is the obvious way to write a coincident population, and it is the wrong one here: it makes what one interferometer's strain contains depend on which other interferometers the run happens to include, so the same channel stops being reproducible between a three-interferometer run and a two-interferometer one. Paired-geometry comparisons need that reproducibility, so an event that lands nowhere simply lands nowhere, and the multiplicity comes out Binomial rather than being imposed.
Attributes:
| Name | Type | Description |
|---|---|---|
participation_probability |
float
|
Probability that any one interferometer the model
applies to receives a given network event. |
amplitude_ratio_std |
float
|
Standard deviation of a per-interferometer lognormal
amplitude ratio with linear mean |
__post_init__()
¶
Validate the configured coherence parameters.
Raises:
| Type | Description |
|---|---|
ValueError
|
If the participation probability is outside |
amplitude_ratio(rng)
¶
Draw one interferometer's amplitude ratio for a shared waveform.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rng
|
Generator
|
The generator for this (event, interferometer) pair. |
required |
Returns:
| Type | Description |
|---|---|
float
|
The multiplier to apply to the shared waveform. |
serialize()
¶
Return the mapping that reconstructs this coherence specification.
NoiseComponentConfig
¶
NoiseConfig
¶
Bases: BaseModel
Generic configuration for composed detector-noise simulations.
OutputConfig
¶
Bases: BaseModel
Configuration for simulation output.
OverlapSaveFirSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate coloured detector noise with a minimum-phase overlap-save FIR.
filter_taps
property
¶
Return a copy of the designed colouring filter taps.
metadata
property
¶
Return simulator metadata.
resume_metadata_nbytes
property
¶
Serialized size of the resume metadata in bytes.
state_nbytes
property
¶
Serialized size of the continuation state in bytes.
target_psd
property
¶
Return a copy of the band-masked target PSD the filter was designed from.
__init__(*, psd_file=None, target_psd=None, target_frequencies=None, filter_length=DEFAULT_FILTER_LENGTH, minimum_phase=True, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE)
¶
Initialize the simulator and design the colouring filter once.
The target may be given either as a file (psd_file) or directly as a
one-sided array (target_psd), so an analytic target can bypass the
tabulated-curve interpolation entirely. When target_frequencies is
omitted for an array target it defaults to the one-sided grid matching
its length at the configured sampling frequency.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_file
|
str | Path | None
|
Path to a two-column PSD table. |
None
|
target_psd
|
ndarray | None
|
One-sided target PSD array, mutually exclusive with
|
None
|
target_frequencies
|
ndarray | None
|
Frequencies of |
None
|
filter_length
|
int
|
Truncation control |
DEFAULT_FILTER_LENGTH
|
minimum_phase
|
bool
|
Whether to apply the cepstral minimum-phase factorisation. |
True
|
detectors
|
list[str] | None
|
Detector names to generate. |
None
|
sampling_frequency
|
float
|
Sampling frequency in hertz. |
4096.0
|
duration
|
float
|
Nominal generation duration in seconds. |
4.0
|
seed
|
int | None
|
Base seed for the per-detector generators. |
None
|
low_frequency_cutoff
|
float
|
Lower edge of the target band in hertz. |
2.0
|
high_frequency_cutoff
|
float | None
|
Upper edge of the target band in hertz. |
None
|
block_size
|
int
|
Overlap-save block length in samples. |
DEFAULT_BLOCK_SIZE
|
continuation_state()
¶
Return the bounded state needed to keep a running stream going.
The state is the filter memory of every detector plus the pending output and block bookkeeping. It contains no generated strain and its size is independent of the generated span.
export_state()
¶
Return a picklable snapshot sufficient to resume the stream.
from_component(component, config)
classmethod
¶
Construct an overlap-save FIR simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector coloured noise with continuity across calls.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield coloured-noise chunks lazily while preserving filter memory.
import_state(state)
¶
Restore a snapshot produced by :meth:export_state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same configuration. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator's settings. |
reset()
¶
Clear the filter memory, pending samples and random-number generators.
resume_metadata()
¶
Return the metadata a stopped stream must persist to resume.
This is the per-detector bit-generator state plus the settings needed to rebuild the filter, and is reported separately from the continuation state.
ParallelAdapter
¶
Wrap an independent-detector simulator factory with parallel execution.
metadata
property
¶
Return metadata describing the wrapped simulator and executor.
segment_gps_start
property
writable
¶
Return the GPS epoch assigned for the next segment, or None if never set.
__init__(base_factory, *, max_workers=None, backend='auto')
¶
Store the simulator factory and executor preferences.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector strain arrays in parallel.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield parallel-generated chunks lazily.
reset()
¶
Clear cached worker simulators, resetting continuity across workers.
PowerLawSNRDistribution
dataclass
¶
Bases: SNRDistribution
Power-law target SNR above a threshold.
With maximum unset the survival function above minimum goes as
S(s) = (s / minimum) ** -alpha, which is the convention a Hill or
maximum-likelihood tail index is quoted in: a class measured to have a tail
index of 1.34 above SNR 10 is configured as
{distribution = "power_law", minimum = 10.0, alpha = 1.34} with no
conversion. Note that alpha is the survival exponent, one less than the
density's. Setting maximum renormalizes that survival onto
[minimum, maximum] as (S(s) - S(maximum)) / (1 - S(maximum)), which
reaches zero at the cap rather than continuing past it.
minimum is a threshold, not a fit to the whole population: the measured
index describes the tail above it and says nothing about the bulk below, so
a model configured this way reproduces the tail and replaces the bulk with
the same power law continued down to the threshold.
maximum truncates the draw. It is optional and unset by default, but for
a heavy tail it is worth setting: with alpha below 1 the untruncated
distribution has no finite mean, and a long enough run will eventually draw
an SNR no detector could produce. The largest SNR observed for the class is
the natural choice.
Below alpha of about 0.052 it stops being optional and is required: the
untruncated draw then runs off the top of the float range (see
:func:_untruncated_headroom_bits), which a configuration cannot express and
a run cannot use, so such a configuration is refused at construction rather
than left to fail partway through a run. Every index in the measured
range -- 0.40 for the heaviest Gravity Spy class, 3.11 for the lightest --
is far above that, and is unaffected.
Attributes:
| Name | Type | Description |
|---|---|---|
minimum |
float
|
Threshold SNR; every draw is at or above it. |
alpha |
float
|
Tail index of the survival function. Must be greater than zero. |
maximum |
float | None
|
Optional upper truncation, strictly above |
distribution |
str
|
Discriminator naming this shape in a configuration. |
__post_init__()
¶
Validate the configured power-law parameters.
sample(rng)
¶
Draw one target SNR by inverting the survival function.
rng.random() returns a uniform on [0, 1), so the inverted
survival minimum * (1 - u) ** (-1 / alpha) covers [minimum, inf)
and never divides by zero. Truncation rescales the same inversion onto
[minimum, maximum].
The untruncated branch is guarded rather than trusted. __post_init__
already refuses a configuration whose draws can leave the float range, so
nothing built through it reaches the guard; what the guard covers is a
value that got past that check anyway -- an instance mutated after
construction, or a configuration sitting on the boundary where the bound
is computed in logarithms and can be a rounding hair out. An unusable
target has to stop here either way, because past this point it is a scale
factor on a waveform and a number in a truth catalogue, where infinity is
indistinguishable from a very loud glitch.
serialize()
¶
Return the mapping that reconstructs this distribution.
survival(threshold)
¶
Return the fraction of draws at or above threshold.
The selection factor of a recording or detection cut applied to this class: how much of the configured population a search at that threshold would ever see. Read off the distribution in closed form rather than estimated from a realization, so an expected-count decomposition is exact rather than Monte-Carlo.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
threshold
|
float
|
The SNR cut. |
required |
Returns:
| Type | Description |
|---|---|
float
|
One at or below |
float
|
and the survival function in between. |
SNRDistribution
dataclass
¶
ScatteredLightGlitch
dataclass
¶
Bases: GlitchModel
Arch-shaped scattered-light transient with a Gaussian envelope.
__post_init__()
¶
Validate scattered-light parameters.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate a chirping scattered-light glitch.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
SchumannNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate correlated strain noise from an isotropic Schumann-resonance model.
metadata
property
¶
Return simulator metadata.
previous_strain
property
¶
Expose continuity buffers for protocol-compatible state inspection.
__init__(*, positions, coupling_files, detectors=None, schumann_params=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, window_duration=DEFAULT_WINDOW_DURATION)
¶
Initialize the Schumann simulator.
from_component(component, config)
classmethod
¶
Construct a Schumann-noise simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate correlated Schumann strain noise.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield Schumann-noise chunks lazily while preserving simulator state.
reset()
¶
Clear continuity and RNG state.
spectrum(frequencies)
¶
Return the isotropic Schumann magnetic PSD on a frequency grid.
theoretical_coherence(frequency, detector_a, detector_b)
¶
Return the isotropic Schumann coherence approximation.
SchumannParams
dataclass
¶
Physical parameters for the isotropic Schumann-resonance model.
__post_init__()
¶
Validate the resonance parameter vectors.
SimulationResult
dataclass
¶
Result of a noise simulation run.
Attributes:
| Name | Type | Description |
|---|---|---|
output_paths |
dict[str, Path]
|
Paths to generated output files, keyed by detector name. |
config |
NoiseConfig
|
The configuration used for the simulation. |
SpectralCovariance
dataclass
¶
Spectral covariance arrays on a common frequency grid.
SpectralLine
dataclass
¶
Configuration for a narrow-band spectral line.
__post_init__()
¶
Validate scalar line parameters.
SpectralLineSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate additive spectral lines directly in the time domain.
metadata
property
¶
Return simulator metadata.
__init__(*, lines, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None)
¶
Initialize the line generator.
from_component(component, config)
classmethod
¶
Construct a spectral-line simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate line-only strain for all requested detectors.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield spectral-line chunks lazily.
reset()
¶
Clear continuity state and phase initialization.
WhiteNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate Gaussian white noise as a composable component.
metadata
property
¶
Return metadata describing the white-noise component.
__init__(*, duration=4.0, sampling_frequency=4096.0, detectors=None, seed=None)
¶
Initialize the white-noise component state.
from_component(component, config)
classmethod
¶
Construct one white-noise component from generic config.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Return Gaussian white-noise strain arrays.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield white-noise strain chunks lazily.
__getattr__(name)
¶
Lazily resolve optional top-level exports.
apply_segment_gps_start(simulator, gps_start)
¶
Point every glitch injector inside a simulator chain at one segment epoch.
A writer that assigns a per-segment epoch -- FrameWriter, whose frame names
carry it -- is the authority on where a segment sits in GPS time, and the glitch
truth catalogue has to be stamped with that same instant or the catalogue and the
frame name describe different data. Threading the writer's epoch through is what
keeps it one number rather than two that agree only while the segments happen to be
contiguous.
Walks the wrapper chain (base) and a composite's components, because the
injector is rarely the outermost simulator: a glitch component sits inside
CompositeNoiseSimulator, which sits inside whatever adapter the writer holds.
A simulator with no injector anywhere beneath it is left alone.
A wrapper that builds its simulators rather than holding one -- ParallelAdapter,
which constructs and caches one per detector from a factory -- cannot be reached by
walking attributes, and an epoch set on such a wrapper alone reached none of its
workers. Those declare a writable segment_gps_start and own the fan-out, so this
hands the epoch over and lets them place both the workers they already hold and any
they build later in the same segment.
assemble_hermitian_spectral_matrices(detectors, psd, csd=None)
¶
Assemble one Hermitian spectral covariance matrix per frequency.
available_glitch_populations()
¶
Return the registered population names, sorted.
available_joint_backend_names()
¶
Return the registered joint backend names, sorted.
build_spectral_covariance_from_files(*, detectors, psd_files, csd_files, frequencies, taper, delta_frequency, regularization_epsilon=1e-12)
¶
Load PSD/CSD files and build spectral covariance products.
cholesky_factors_from_spectral_matrices(spectral_matrices, *, delta_frequency, regularization_epsilon=1e-12)
¶
Build coefficient-covariance Cholesky factors from one-sided spectra.
compare_psd(data, target_psd_file, sampling_frequency, rtol=0.1, fmin=10.0, fmax=None)
¶
Compare an estimated PSD against a target PSD file over a frequency band.
This helper compares the median PSD level in-band rather than every bin individually, which makes it robust to the residual variance of a single Welch estimate from stochastic noise.
discover_joint_backends()
cached
¶
Return the entry points registered under :data:JOINT_BACKEND_ENTRY_POINT_GROUP.
This performs discovery only -- it does not import any backend's module.
Call :meth:importlib.metadata.EntryPoint.load on a returned entry point
(or use :func:load_joint_backend) to import and resolve the backend
class it names.
estimate_psd(data, sampling_frequency, segment_duration=4.0, overlap=0.5)
¶
Estimate a one-sided PSD in physical density units.
The returned frequencies are in Hz. For strain input data, the PSD is in strain^2/Hz.
et_o3_anchored_population(*, psd_file=DEFAULT_PSD, participation_probability=COHERENT_PARTICIPATION_PROBABILITY, amplitude_ratio_std=COHERENT_AMPLITUDE_RATIO_STD, coherent_rate_hz=COHERENT_RATE_HZ, short_rate_hz=SHORT_RATE_HZ)
¶
Build the et-o3-anchored-v1 population.
Called with no arguments this returns the registered population, and its digest is the one a campaign pins. The arguments exist so the assumptions that could not be anchored can be varied deliberately -- a sweep over the participation probability is the obvious one -- and any such variation produces a different digest, which is the point: a run cannot claim to have used the registered population while having weakened it.
Note that short_rate_hz and coherent_rate_hz are the rates of two different
kinds of process. The short class runs one Poisson process per interferometer, so
its rate is what each of them sees. The longer class runs one process for the whole
network, so its rate is the rate of shared events; each interferometer sees that times
the participation probability.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_file
|
str | Path
|
Noise curve the target SNRs are calibrated against. A bundled preset name keeps the digest portable; an absolute path does not. |
DEFAULT_PSD
|
participation_probability
|
float
|
Probability that one interferometer receives a given
event of the coherent class. Unanchored; see :data: |
COHERENT_PARTICIPATION_PROBABILITY
|
amplitude_ratio_std
|
float
|
Spread of the per-interferometer amplitude ratio of the coherent class. Unanchored. |
COHERENT_AMPLITUDE_RATIO_STD
|
coherent_rate_hz
|
float
|
Network event rate of the coherent class, in Hz. |
COHERENT_RATE_HZ
|
short_rate_hz
|
float
|
Per-interferometer event rate of the short class, in Hz. |
SHORT_RATE_HZ
|
Returns:
| Type | Description |
|---|---|
GlitchPopulation
|
The population, with both classes scoped to every interferometer in the run. |
get_glitch_population(name)
¶
Return a registered population by name.
Each call builds a fresh population, so a caller cannot mutate the registry's copy out from under the next caller. Two calls compare equal and digest identically.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The registered name, e.g. |
required |
Returns:
| Type | Description |
|---|---|
GlitchPopulation
|
The population. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If no population is registered under that name. The message lists the
names that are, because a typo is the likeliest cause and a bare |
interpolate_complex_spectral_series(frequencies, values, target_frequencies, *, taper=None)
¶
Interpolate a complex spectral series by real and imaginary parts.
interpolate_real_spectral_series(frequencies, values, target_frequencies, *, taper=None)
¶
Interpolate and clip a real non-negative spectral series.
load_and_interpolate_csd(file_path, target_frequencies, *, taper=None)
¶
Load and interpolate a complex one-sided CSD file.
load_and_interpolate_psd(file_path, target_frequencies, *, taper=None)
¶
Load and interpolate a one-sided PSD file.
load_config(path)
¶
Load and validate noise configuration from a file.
Supports TOML, YAML, and JSON formats.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path | str
|
Path to the configuration file. |
required |
Returns:
| Type | Description |
|---|---|
NoiseConfig
|
Validated NoiseConfig instance. |
Raises:
| Type | Description |
|---|---|
FileNotFoundError
|
If the config file does not exist. |
ValueError
|
If the file format is unsupported or validation fails. |
load_joint_backend(name)
¶
Import and return the backend class registered under name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The entry-point name a package registered its backend under. |
required |
Returns:
| Type | Description |
|---|---|
Any
|
The backend class (or factory callable) the entry point names. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If no backend is registered under |
normalize_csd_mapping(csd_values, *, detectors)
¶
Validate and normalize detector-pair CSD mappings.
normalize_detector_pair(pair)
¶
Normalize a detector-pair key to sorted order.
open_stream(simulator, *, chunk_duration, sampling_frequency, detectors, seed=None)
¶
Open a stateful chunk stream on a protocol-compatible simulator.
regularized_cholesky(spectral_matrix, *, regularization_epsilon=1e-12)
¶
Return a numerically stable Cholesky factor for a Hermitian matrix.
run_diagnostics(data, sampling_frequency, whiten=True)
¶
Run the default Gaussianity and stationarity diagnostics.
Data is whitened against its own Welch PSD estimate before testing (see
_whiten_samples), which makes both checks applicable to coloured
noise. Pass whiten=False to test the raw samples instead — only
meaningful for broadband data.
sample_complex_frequency_coefficients(rng, cholesky_factors, *, real_only_indices=None)
¶
Sample complex Gaussian frequency coefficients from Cholesky factors.
simulate_spectral_covariance_chunk(rng, cholesky_factors, *, detectors, frequency_grid_size, frequency_mask, delta_frequency, window_size)
¶
Sample one real multi-detector chunk from spectral covariance factors.
take(stream, total_duration, chunk_duration, sampling_frequency)
¶
Collect stream chunks up to total_duration seconds.
time_series_from_frequency_coefficients(coefficients, *, detectors, frequency_grid_size, frequency_mask, delta_frequency, window_size)
¶
Convert masked detector frequency coefficients to real time series.
Configuration¶
Pydantic schema models and load_config for TOML/YAML/JSON.
Configuration schema and loading helpers for gwmock_noise.
NoiseComponentConfig
¶
NoiseConfig
¶
Bases: BaseModel
Generic configuration for composed detector-noise simulations.
OutputConfig
¶
Bases: BaseModel
Configuration for simulation output.
load_config(path)
¶
Load and validate noise configuration from a file.
Supports TOML, YAML, and JSON formats.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path | str
|
Path to the configuration file. |
required |
Returns:
| Type | Description |
|---|---|
NoiseConfig
|
Validated NoiseConfig instance. |
Raises:
| Type | Description |
|---|---|
FileNotFoundError
|
If the config file does not exist. |
ValueError
|
If the file format is unsupported or validation fails. |
Gaussian noise models and helpers¶
Spectral-line definitions plus PSD reference helpers used by Gaussian-noise configuration and simulators.
Gaussian-noise domain models and helpers.
SpectralLine
dataclass
¶
Configuration for a narrow-band spectral line.
__post_init__()
¶
Validate scalar line parameters.
is_remote_psd_reference(value)
¶
Return whether a PSD reference is an HTTP(S) URL.
normalize_spectral_lines(value)
¶
Normalize heterogeneous spectral-line inputs to dataclass instances.
resolve_bundled_psd_preset(value)
¶
Resolve a bare preset name to a bundled PSD asset.
Spectral covariance utilities¶
Signal-agnostic helpers for loading PSD/CSD files, interpolating spectra onto an FFT grid, assembling Hermitian covariance matrices, regularizing Cholesky factors, and sampling real multi-detector time series.
Signal-agnostic spectral covariance utilities.
The utilities in this module use one-sided PSD/CSD inputs in units of strain^2/Hz
and sample masked real-FFT coefficients with covariance S(f) / (2 df). The
inverse transform multiplies by df * n so the generated real time series is
consistent with the one-sided periodogram convention used by the simulators.
SpectralCovariance
dataclass
¶
Spectral covariance arrays on a common frequency grid.
assemble_hermitian_spectral_matrices(detectors, psd, csd=None)
¶
Assemble one Hermitian spectral covariance matrix per frequency.
build_spectral_covariance_from_files(*, detectors, psd_files, csd_files, frequencies, taper, delta_frequency, regularization_epsilon=1e-12)
¶
Load PSD/CSD files and build spectral covariance products.
cholesky_factors_from_spectral_matrices(spectral_matrices, *, delta_frequency, regularization_epsilon=1e-12)
¶
Build coefficient-covariance Cholesky factors from one-sided spectra.
interpolate_complex_spectral_series(frequencies, values, target_frequencies, *, taper=None)
¶
Interpolate a complex spectral series by real and imaginary parts.
interpolate_real_spectral_series(frequencies, values, target_frequencies, *, taper=None)
¶
Interpolate and clip a real non-negative spectral series.
load_and_interpolate_csd(file_path, target_frequencies, *, taper=None)
¶
Load and interpolate a complex one-sided CSD file.
load_and_interpolate_psd(file_path, target_frequencies, *, taper=None)
¶
Load and interpolate a one-sided PSD file.
normalize_csd_mapping(csd_values, *, detectors)
¶
Validate and normalize detector-pair CSD mappings.
normalize_detector_pair(pair)
¶
Normalize a detector-pair key to sorted order.
regularized_cholesky(spectral_matrix, *, regularization_epsilon=1e-12)
¶
Return a numerically stable Cholesky factor for a Hermitian matrix.
sample_complex_frequency_coefficients(rng, cholesky_factors, *, real_only_indices=None)
¶
Sample complex Gaussian frequency coefficients from Cholesky factors.
simulate_spectral_covariance_chunk(rng, cholesky_factors, *, detectors, frequency_grid_size, frequency_mask, delta_frequency, window_size)
¶
Sample one real multi-detector chunk from spectral covariance factors.
time_series_from_frequency_coefficients(coefficients, *, detectors, frequency_grid_size, frequency_mask, delta_frequency, window_size)
¶
Convert masked detector frequency coefficients to real time series.
Glitch models¶
Runtime glitch model definitions, including phenomenological models, the optional gengli-backed blip implementation, and the optional DeepExtractor model that injects real O3 glitch reconstructions.
Public glitch-model implementations.
BlipGlitch
dataclass
¶
Bases: GlitchModel
Gaussian-windowed broadband burst.
__post_init__()
¶
Validate blip-specific parameters.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate a Gaussian-windowed white-noise burst.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
DeepExtractorGlitch
dataclass
¶
Bases: GlitchModel
Real O3 glitch reconstructions colored against a target PSD.
Waveforms come from the DeepExtractor glitch-reconstruction dataset:
whitened, amplitude-normalized 2-second time series sampled at 4096 Hz,
covering seven Gravity Spy classes. Each drawn waveform is resampled to
the simulation rate, colored against the configured PSD, and rescaled so
its optimal SNR sqrt(4 df sum(|h(f)|^2 / S(f))) against that PSD
matches the configured target. The dataset (~2.3 GB) is downloaded lazily
on first use and cached by huggingface_hub; later runs reuse the cached
files. On each run the Hub is contacted first to validate the cached
files' ETags and download anything missing. If the Hub is unreachable, a
warning notes that the ETag check was skipped and the cached files are used
instead; if they are not cached, the first generate_waveform raises
huggingface_hub's LocalEntryNotFoundError (a FileNotFoundError
subclass). Set local_files_only=True to skip the network unconditionally
and read straight from the cache.
revision pins the download to a specific dataset version — a git branch,
tag, or commit SHA forwarded to hf_hub_download. Leave it None to
track the repository default. Either way, the first download resolves to a
concrete commit SHA (read back from the Hub cache path), which is what
serialize records. Replaying a run through its metadata therefore fetches
the exact commit that produced it, keeping glitch generation bit-reproducible
for a fixed (version, config, seed) even as the upstream dataset moves.
rate accepts either a single number — the total Poisson rate shared by
all configured classes, drawn uniformly — or a mapping from class name to
per-class Poisson rate, in which case the total rate is their sum and each
event's class is drawn proportionally to its rate.
snr accepts a number, a mapping from class name to number, a distribution,
or a mapping from class name to distribution — freely mixed, so one class can
be sampled while another stays fixed. A number calibrates every event of that
class to exactly that loudness; a distribution draws a target per event from
the model's own random stream, which is what a measured glitch population
needs, since each Gravity Spy class is heavy-tailed and the tail indices
differ by close to an order of magnitude between classes. See
:mod:gwmock_noise.glitches.snr for the supported shapes.
What the target means is the caller's to decide. A tail index measured
from a trigger generator's SNR — Omicron's, say — is defined against that
pipeline's PSD over the band it searched, while this model calibrates against
psd_file from low_frequency_cutoff upward. Configuring one from the
other equates two SNRs defined against different noise curves over different
bands. The sampler reproduces whatever distribution it is given and cannot
tell whether that identification is the one you meant.
Resampling uses linear interpolation without an anti-aliasing filter, so sampling frequencies below 4096 Hz alias high-frequency content; the SNR calibration itself is unaffected because it is computed after resampling.
__post_init__()
¶
Validate the configuration and preload the PSD table.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate one colored, SNR-calibrated DeepExtractor glitch.
resolve()
¶
Pin the dataset to a concrete commit SHA, fetching only the samples file.
Reads the commit SHA huggingface_hub stamps into the cache path so the
resolved revision is available for reproducibility metadata even when no
waveform has been drawn yet — e.g. a batch that fires zero glitch events,
where generate_waveform (and therefore _get_dataset) never runs.
The samples file is downloaded once and reused by the later full load, so
this adds no extra download on the generation path. Returns the resolved
SHA, or None when the cache layout does not expose one (e.g. a test
stub serving files from a flat directory).
serialize()
¶
Return metadata-friendly model parameters.
EmpiricalSNRDistribution
dataclass
¶
Bases: SNRDistribution
Target SNR drawn with replacement from a table of observed SNRs.
For a caller holding the measurements rather than a fit to them: the drawn
population is exactly the supplied one, tail included, with no assumption
about its shape. Supply the values inline as samples or point file
at a table on disk -- exactly one of the two.
Attributes:
| Name | Type | Description |
|---|---|---|
samples |
Sequence[float] | None
|
Observed SNRs given inline, or |
file |
str | Path | None
|
Path to a table of observed SNRs, or |
distribution |
str
|
Discriminator naming this shape in a configuration. |
__post_init__()
¶
Validate the configuration and load the SNR table.
sample(rng)
¶
Draw one target SNR uniformly from the table, with replacement.
serialize()
¶
Return the mapping that reconstructs this distribution.
Echoes whichever of the two inputs was configured, so a file-backed table replays as a reference to the file rather than as a copy of its contents inlined into the metadata.
survival(threshold)
¶
Return the fraction of the supplied table at or above threshold.
GengliBlipGlitch
dataclass
¶
Bases: GlitchModel
File-backed gengli blip generator colored against a target PSD.
__post_init__()
¶
Validate configured files and preload the population/PSD tables.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
from_population_file(population_file, **kwargs)
classmethod
¶
Construct a gengli glitch model from a population file path.
generate_waveform(sampling_frequency, rng=None)
¶
Generate one colored gengli blip waveform.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
GlitchDraw
¶
Bases: NamedTuple
One drawn glitch waveform and the parameters that produced it.
What :meth:GlitchModel.draw returns so an injector can record what it
injected, not merely that it injected something. Every field but the
waveform is optional because not every model has the quantity: a parametric
blip with no PSD has no SNR to report, and only the dataset-backed models
draw a Gravity Spy class.
Attributes:
| Name | Type | Description |
|---|---|---|
waveform |
ndarray
|
The glitch strain, sampled at the requested rate. Its first sample lands at the injection time -- the waveform starts there, it does not peak there. |
amplitude |
float | None
|
The amplitude multiplier drawn from the model's amplitude
distribution, or |
glitch_class |
str | None
|
The drawn morphological class (e.g. a Gravity Spy class
name), or |
target_snr |
float | None
|
The optimal SNR the draw was calibrated to before the
amplitude multiplier, or |
realized_snr |
float | None
|
The optimal SNR of the whole of How it relates to |
GlitchModel
dataclass
¶
Base dataclass for transient glitch generators.
detectors scopes the model to a subset of the run's interferometers. It
accepts one name or a list of them, and defaults to None, which is every
interferometer -- what a model without a selector has always meant. Scoping is
what lets one configuration carry a 10 km model for one set of interferometers
and a 15 km model for another, and per-interferometer rates for a network whose
instruments are not equally glitchy (in O3, Fast_Scattering fired about 29 times
more often in L1 than in H1).
A selector is only half of it. :func:validate_glitch_detector_coverage refuses
a run whose models do not between them account for every interferometer, or
whose models color one interferometer against two different PSDs -- see its
docstring for why proceeding was the defect.
__post_init__()
¶
Validate common glitch parameters.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate a single glitch waveform.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
LogNormalAmplitudeDistribution
dataclass
¶
NetworkCoherence
dataclass
¶
Share one glitch model's events across the interferometers it applies to.
Without it, every glitch model runs an independent Poisson process and an independent random stream per interferometer, so two interferometers never carry the same transient. That is the right default, and it is what the one measurement of the question says for widely separated sites: over O1 and O2 the LIGO blip population produced no coincidences inside the ±15 ms window an astrophysical signal can occupy, and the coincidences found in wider windows matched the accidental expectation.
It is not what a network of co-located interferometers is expected to do. Three interferometers sharing one site, one vacuum system and one seismic environment see a common environmental transient in all three at once, and a coincidence veto -- or a null stream -- assumes exactly the incoherence that such an event breaks. A model carrying this runs one Poisson process for the whole network, draws one waveform per event, and offers that waveform to each interferometer it applies to.
Participation is decided per interferometer, independently, and is never conditioned on how many others took the event. Conditioning -- "keep only events that land in at least two interferometers" -- is the obvious way to write a coincident population, and it is the wrong one here: it makes what one interferometer's strain contains depend on which other interferometers the run happens to include, so the same channel stops being reproducible between a three-interferometer run and a two-interferometer one. Paired-geometry comparisons need that reproducibility, so an event that lands nowhere simply lands nowhere, and the multiplicity comes out Binomial rather than being imposed.
Attributes:
| Name | Type | Description |
|---|---|---|
participation_probability |
float
|
Probability that any one interferometer the model
applies to receives a given network event. |
amplitude_ratio_std |
float
|
Standard deviation of a per-interferometer lognormal
amplitude ratio with linear mean |
__post_init__()
¶
Validate the configured coherence parameters.
Raises:
| Type | Description |
|---|---|
ValueError
|
If the participation probability is outside |
amplitude_ratio(rng)
¶
Draw one interferometer's amplitude ratio for a shared waveform.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rng
|
Generator
|
The generator for this (event, interferometer) pair. |
required |
Returns:
| Type | Description |
|---|---|
float
|
The multiplier to apply to the shared waveform. |
serialize()
¶
Return the mapping that reconstructs this coherence specification.
PowerLawSNRDistribution
dataclass
¶
Bases: SNRDistribution
Power-law target SNR above a threshold.
With maximum unset the survival function above minimum goes as
S(s) = (s / minimum) ** -alpha, which is the convention a Hill or
maximum-likelihood tail index is quoted in: a class measured to have a tail
index of 1.34 above SNR 10 is configured as
{distribution = "power_law", minimum = 10.0, alpha = 1.34} with no
conversion. Note that alpha is the survival exponent, one less than the
density's. Setting maximum renormalizes that survival onto
[minimum, maximum] as (S(s) - S(maximum)) / (1 - S(maximum)), which
reaches zero at the cap rather than continuing past it.
minimum is a threshold, not a fit to the whole population: the measured
index describes the tail above it and says nothing about the bulk below, so
a model configured this way reproduces the tail and replaces the bulk with
the same power law continued down to the threshold.
maximum truncates the draw. It is optional and unset by default, but for
a heavy tail it is worth setting: with alpha below 1 the untruncated
distribution has no finite mean, and a long enough run will eventually draw
an SNR no detector could produce. The largest SNR observed for the class is
the natural choice.
Below alpha of about 0.052 it stops being optional and is required: the
untruncated draw then runs off the top of the float range (see
:func:_untruncated_headroom_bits), which a configuration cannot express and
a run cannot use, so such a configuration is refused at construction rather
than left to fail partway through a run. Every index in the measured
range -- 0.40 for the heaviest Gravity Spy class, 3.11 for the lightest --
is far above that, and is unaffected.
Attributes:
| Name | Type | Description |
|---|---|---|
minimum |
float
|
Threshold SNR; every draw is at or above it. |
alpha |
float
|
Tail index of the survival function. Must be greater than zero. |
maximum |
float | None
|
Optional upper truncation, strictly above |
distribution |
str
|
Discriminator naming this shape in a configuration. |
__post_init__()
¶
Validate the configured power-law parameters.
sample(rng)
¶
Draw one target SNR by inverting the survival function.
rng.random() returns a uniform on [0, 1), so the inverted
survival minimum * (1 - u) ** (-1 / alpha) covers [minimum, inf)
and never divides by zero. Truncation rescales the same inversion onto
[minimum, maximum].
The untruncated branch is guarded rather than trusted. __post_init__
already refuses a configuration whose draws can leave the float range, so
nothing built through it reaches the guard; what the guard covers is a
value that got past that check anyway -- an instance mutated after
construction, or a configuration sitting on the boundary where the bound
is computed in logarithms and can be a rounding hair out. An unusable
target has to stop here either way, because past this point it is a scale
factor on a waveform and a number in a truth catalogue, where infinity is
indistinguishable from a very loud glitch.
serialize()
¶
Return the mapping that reconstructs this distribution.
survival(threshold)
¶
Return the fraction of draws at or above threshold.
The selection factor of a recording or detection cut applied to this class: how much of the configured population a search at that threshold would ever see. Read off the distribution in closed form rather than estimated from a realization, so an expected-count decomposition is exact rather than Monte-Carlo.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
threshold
|
float
|
The SNR cut. |
required |
Returns:
| Type | Description |
|---|---|
float
|
One at or below |
float
|
and the survival function in between. |
SNRDistribution
dataclass
¶
ScatteredLightGlitch
dataclass
¶
Bases: GlitchModel
Arch-shaped scattered-light transient with a Gaussian envelope.
__post_init__()
¶
Validate scattered-light parameters.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate a chirping scattered-light glitch.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
load_snr_samples(path)
¶
Load a one-dimensional table of observed SNRs from disk.
.h5/.hdf5 files are read from their snr dataset -- the schema
gwmock-noise build-blip-glitch-table writes -- and anything else is read
as a whitespace-separated text column.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str | Path
|
Path to the SNR table. |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
The validated SNR samples. |
Raises:
| Type | Description |
|---|---|
FileNotFoundError
|
If |
ValueError
|
If the table is empty, not one-dimensional, or holds a non-finite or non-positive SNR. |
normalize_glitch_models(value)
¶
Normalize heterogeneous glitch-config inputs.
Entries carry an optional detectors selector -- one interferometer name or a
list of them -- which scopes the model to those interferometers; an entry
without one applies to every interferometer in the run, as it always has.
Whether the resulting list accounts for the run is
:func:validate_glitch_detector_coverage's question, because it is the run that
supplies the interferometers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
Any
|
The configured list of models or model mappings. |
required |
Returns:
| Type | Description |
|---|---|
list[GlitchModel]
|
The normalized glitch models. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If the value is not a list, or an entry is not a mapping. |
ValueError
|
If an entry names an unsupported kind or omits its amplitude distribution. |
normalize_snr(value, parameter='snr')
¶
Normalize one target-SNR specification into a number or a distribution.
Accepts what a configuration can hold for a single target: a number, a
already-constructed :class:SNRDistribution, or a mapping describing one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
Any
|
The configured specification. |
required |
parameter
|
str
|
Name of the configuration key, used in error messages. |
'snr'
|
Returns:
| Type | Description |
|---|---|
float | SNRDistribution
|
The target SNR as a float, or the distribution to draw it from. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
ValueError
|
If a numeric target is not finite and positive. |
snr_survival(specification, threshold)
¶
Return the fraction of one class's events whose target SNR reaches threshold.
The selection factor in an expected-count decomposition. A model with no target SNR at
all has nothing to select on, and reports 1.0 rather than guessing: its loudness is
set by the amplitude distribution in units the threshold is not expressed in.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
specification
|
float | SNRDistribution | None
|
A fixed target SNR, a distribution to draw it from, or |
required |
threshold
|
float
|
The SNR cut. |
required |
Returns:
| Type | Description |
|---|---|
float
|
The surviving fraction, between zero and one. |
supported_glitch_kinds()
¶
Return all supported glitch-model kinds.
validate_glitch_detector_coverage(models, detectors)
¶
Refuse a glitch configuration that does not say what every interferometer gets.
A glitch model used to apply to every interferometer in the run because there
was no way to say otherwise, so one psd_file colored all of them. A run
naming both Einstein Telescope designs resolves to interferometers of two arm
lengths, and the 10 km curve was applied to the 15 km instruments as well: SNRs
23-44% away from what the configuration asked for, morphology-dependent, with
nothing in the output or the logs to say so. The silence was the defect, so
a configuration that cannot be read one way is refused here rather than run.
Three things are refused, in the order a reader can act on them:
- A selector naming an interferometer the run does not have. Usually a typo or a leftover from another network, and the model then never fires -- a configured rate and PSD that inject nothing, which nothing in the output distinguishes from a model that fired and drew no events.
- An interferometer no model claims. Its strain is written with no glitches
while its neighbours carry them. Where that is the intent, say it: a model
scoped to it with
rate = 0.0states "this interferometer carries no glitches" in the configuration, where a reader can see it. - Two coloring PSDs claiming one interferometer. An interferometer has one noise floor; two models coloring it against different curves disagree about what instrument it is, and whichever ran second would not make that true. Models that do no coloring impose no floor and are not part of this.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
models
|
Sequence[GlitchModel]
|
The run's glitch models, already normalized. |
required |
detectors
|
Sequence[str]
|
The interferometers the run will generate. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If any of the three cases above holds. The message names the interferometers involved. |
Glitch populations¶
Registered glitch populations: a named, serializable, digestible set of glitch classes, with an expected-count decomposition and a reusable seeded realization that paired arms of a comparison share.
Registered glitch populations and the machinery that pins them.
ExpectedCount
dataclass
¶
The expected injected count for one class in one interferometer, factorized.
The product is :attr:expected, and every factor is carried beside it. A campaign that
finds the total surprising can then read off which factor is surprising -- a rate
transplanted from another run, a livetime that is calendar time rather than observing
time, a participation probability nobody meant to set -- instead of re-deriving the
whole chain.
Attributes:
| Name | Type | Description |
|---|---|---|
glitch_class |
str
|
Name of the class within its population. |
detector |
str
|
The interferometer this row is for. |
rate_hz |
float
|
The class's Poisson rate. For a class whose interferometers glitch independently this is the rate each of them sees; for a network-coherent class it is the rate of the shared network process, which is not the same number unless every interferometer takes every event. |
livetime_seconds |
float
|
The analysed time this count is over. |
participation |
float
|
Probability that this interferometer receives a given event of the class. One for a class with no network coherence. |
selection |
float
|
Fraction of the class's events that survive the SNR cut the count was asked for. One when no cut was given. |
expected |
float
|
|
network_events
property
¶
Expected number of network events of this class over the same livetime.
The same arithmetic with the participation factor divided out: the events the class's process produced, rather than the subset this interferometer received. For a class with no network coherence the two coincide, because every event it draws for an interferometer is that interferometer's.
GlitchClass
dataclass
¶
One named class within a registered population.
The prose fields are part of the registered definition and therefore part of the digest, deliberately. A population whose numbers are unchanged but whose meaning has been rewritten is a different population, and a digest that did not move would say otherwise.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Stable identifier of the class within its population. |
morphology |
str
|
Short slug naming the waveform family, e.g.
|
description |
str
|
One sentence saying what the class represents and what it is drawn from. |
model |
GlitchModel
|
The glitch model that generates it, already normalized. |
__post_init__()
¶
Validate the class definition.
Raises:
| Type | Description |
|---|---|
ValueError
|
If the name, morphology or description is blank. |
deserialize(payload)
classmethod
¶
Rebuild a class from :meth:serialize output.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
payload
|
dict[str, Any]
|
The serialized class. |
required |
Returns:
| Type | Description |
|---|---|
GlitchClass
|
The reconstructed class. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If a required key is missing. |
serialize()
¶
Return the mapping that reconstructs this class.
GlitchPopulation
dataclass
¶
A named, versioned, digestible set of glitch classes.
The digest is only as portable as what goes into it. A model's psd_file is
serialized as written, so a population built against /scratch/me/et.txt carries
that path into its digest and no other machine can reproduce it. Name a bundled preset,
or a path relative to the campaign's own tree, and the digest travels.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Stable registered identifier, e.g. |
description |
str
|
One paragraph saying what the population is for and what it is anchored to. |
classes |
tuple[GlitchClass, ...]
|
The classes, in registration order. |
references |
tuple[str, ...]
|
Published sources the registered quantities are anchored to, one citation per entry. |
unanchored |
tuple[str, ...]
|
Quantities that could not be anchored to a published measurement, one per entry, each saying what it is and why. Carried in the population rather than only in its documentation, so a consumer reading a pinned population reads the caveats with it. An empty tuple is a claim that everything is anchored, so it is a claim to be able to defend. |
schema_version |
str
|
Serialization schema, part of the digest. |
__post_init__()
¶
Validate the population definition.
Raises:
| Type | Description |
|---|---|
ValueError
|
If the name or description is blank, the population has no classes, or two classes share a name. |
build_injector(base, *, gps_start=0.0)
¶
Wrap a base simulator so it also injects this population.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
base
|
NoiseSimulator
|
The simulator producing the noise the glitches are added to. |
required |
gps_start
|
float
|
GPS time of the first sample, stamped onto the truth catalogue. |
0.0
|
Returns:
| Type | Description |
|---|---|
InjectGlitches
|
The wrapped simulator. |
canonical_json()
¶
Return the exact bytes the digest is taken over.
Sorted keys, no insignificant whitespace, ASCII only and no NaN: a mapping that
two processes can be relied on to render identically. Exposed rather than hidden so
a campaign that disagrees about a digest can diff what was hashed instead of
guessing.
Returns:
| Type | Description |
|---|---|
str
|
The canonical serialization. |
deserialize(payload)
classmethod
¶
Rebuild a population from :meth:serialize output.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
payload
|
dict[str, Any]
|
The serialized population. |
required |
Returns:
| Type | Description |
|---|---|
GlitchPopulation
|
The reconstructed population, which serializes and digests identically. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If a required key is missing. |
ValueError
|
If the payload declares a schema version this package cannot read. |
digest()
¶
Return the population's sha256:... digest.
What a campaign manifest pins. Every registered quantity, every prose field and the schema version go into it, so a population that has been edited cannot present itself as the one a finished run used.
Returns:
| Type | Description |
|---|---|
str
|
The digest, algorithm-prefixed. |
expected_counts(*, livetime_seconds, detectors, snr_threshold=None)
¶
Decompose the expected injected count per class and interferometer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
livetime_seconds
|
float
|
Analysed time the count is over. This is observing time, not calendar time: the distinction is the commonest factor-of-1.3 error in a count derived from a published rate. |
required |
detectors
|
Sequence[str]
|
The run's interferometers. Checked against the population's own scoping, so a decomposition cannot be computed for a network the population does not account for. |
required |
snr_threshold
|
float | None
|
Optional recording or detection cut. When given, the selection factor is the fraction of each class's target-SNR distribution at or above it, read off the distribution in closed form. |
None
|
Returns:
| Type | Description |
|---|---|
list[ExpectedCount]
|
One row per (class, interferometer) pair, in registration order. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the livetime is not positive, no interferometers were given, or the population does not account for every interferometer. |
model_for(glitch_class)
¶
Return one class's glitch model by name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
glitch_class
|
str
|
The class name. |
required |
Returns:
| Type | Description |
|---|---|
GlitchModel
|
The model that generates it. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If the population has no such class. |
models()
¶
Return the population's glitch models in registration order.
realize(*, detectors, duration, sampling_frequency, seed, gps_start=0.0)
¶
Generate this population's glitch strain on its own, with no noise under it.
This is the object paired arms share. A comparison whose arms differ in their
signal content, their geometry, their detection threshold or their ranking
statistic must not differ in their glitches, and the way to guarantee that is for
every arm to add the same array rather than to re-run a generator and trust it to
agree. The returned stamp carries the population digest and the seed, so an arm can
record which realization it added -- and the digest is the field to pin, for the
reason :class:GlitchRealization gives.
The realization is reproducible for a fixed (package version, population, seed) and does not depend on the base noise, on the other interferometers in the run, or on whether it was generated in one call or streamed -- see the module's tests, which assert each of those separately.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
Sequence[str]
|
The interferometers to realize. |
required |
duration
|
float
|
Length of the realization in seconds. |
required |
sampling_frequency
|
float
|
Sampling frequency in Hz. Note that a class with a
|
required |
seed
|
int
|
The realization's seed. |
required |
gps_start
|
float
|
GPS time of the first sample. |
0.0
|
Returns:
| Type | Description |
|---|---|
GlitchRealization
|
The glitch strain, its truth catalogue and the stamp identifying it. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the population does not account for every interferometer. |
serialize()
¶
Return the mapping that reconstructs this population.
GlitchRealization
¶
Bases: NamedTuple
One seeded glitch realization and the stamp that identifies it.
Attributes:
| Name | Type | Description |
|---|---|---|
strain |
dict[str, ndarray]
|
The glitch contribution per interferometer, with nothing else in it. This is what paired arms add to their own base noise, and adding the same array to two arms is what makes their glitch content bit-identical. |
events |
list[dict[str, Any]]
|
The per-event truth catalogue, in time order. See
:data: |
stamp |
dict[str, Any]
|
What identifies this realization: the population's name and digest, the
package version, and the generation parameters. Enough to reproduce The pinned value is |
available_glitch_populations()
¶
Return the registered population names, sorted.
et_o3_anchored_population(*, psd_file=DEFAULT_PSD, participation_probability=COHERENT_PARTICIPATION_PROBABILITY, amplitude_ratio_std=COHERENT_AMPLITUDE_RATIO_STD, coherent_rate_hz=COHERENT_RATE_HZ, short_rate_hz=SHORT_RATE_HZ)
¶
Build the et-o3-anchored-v1 population.
Called with no arguments this returns the registered population, and its digest is the one a campaign pins. The arguments exist so the assumptions that could not be anchored can be varied deliberately -- a sweep over the participation probability is the obvious one -- and any such variation produces a different digest, which is the point: a run cannot claim to have used the registered population while having weakened it.
Note that short_rate_hz and coherent_rate_hz are the rates of two different
kinds of process. The short class runs one Poisson process per interferometer, so
its rate is what each of them sees. The longer class runs one process for the whole
network, so its rate is the rate of shared events; each interferometer sees that times
the participation probability.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_file
|
str | Path
|
Noise curve the target SNRs are calibrated against. A bundled preset name keeps the digest portable; an absolute path does not. |
DEFAULT_PSD
|
participation_probability
|
float
|
Probability that one interferometer receives a given
event of the coherent class. Unanchored; see :data: |
COHERENT_PARTICIPATION_PROBABILITY
|
amplitude_ratio_std
|
float
|
Spread of the per-interferometer amplitude ratio of the coherent class. Unanchored. |
COHERENT_AMPLITUDE_RATIO_STD
|
coherent_rate_hz
|
float
|
Network event rate of the coherent class, in Hz. |
COHERENT_RATE_HZ
|
short_rate_hz
|
float
|
Per-interferometer event rate of the short class, in Hz. |
SHORT_RATE_HZ
|
Returns:
| Type | Description |
|---|---|
GlitchPopulation
|
The population, with both classes scoped to every interferometer in the run. |
get_glitch_population(name)
¶
Return a registered population by name.
Each call builds a fresh population, so a caller cannot mutate the registry's copy out from under the next caller. Two calls compare equal and digest identically.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The registered name, e.g. |
required |
Returns:
| Type | Description |
|---|---|
GlitchPopulation
|
The population. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If no population is registered under that name. The message lists the
names that are, because a typo is the likeliest cause and a bare |
Simulators¶
Simulator implementations, the NoiseSimulator protocol, the optional
JointStrainWitnessSimulator protocol and its public
JointDummyCorrelatedSimulator backend, ConfigurableNoiseSimulator,
SimulationResult, and streaming helpers such as open_stream and take.
Noise simulators for gravitational wave detectors.
ARMANoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate stationary detector noise with a fitted ARMA / state-space model.
denominator_coefficients
property
¶
Return a copy of the ARMA denominator coefficients a_1 .. a_p.
metadata
property
¶
Return metadata describing the fitted ARMA model.
model_psd_curve
property
¶
Return the fitted model PSD on the band-masked design grid.
numerator_coefficients
property
¶
Return a copy of the ARMA numerator coefficients b_0 .. b_q.
resume_metadata_nbytes
property
¶
Serialized size of the resume metadata in bytes.
state_nbytes
property
¶
Serialized size of the continuation state in bytes.
target_psd_curve
property
¶
Return the band-masked target PSD on the design grid.
__init__(*, psd_file=None, target_psd=None, target_frequencies=None, ar_order=DEFAULT_AR_ORDER, ma_order=DEFAULT_MA_ORDER, line_frequencies=None, line_widths=None, detect_lines=False, max_lines=DEFAULT_MAX_LINES, line_prominence=DEFAULT_LINE_PROMINENCE, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization=DEFAULT_REGULARIZATION)
¶
Initialize the simulator and fit the ARMA model once.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_file
|
str | Path | None
|
Path or bundled preset name of a two-column PSD table. |
None
|
target_psd
|
ndarray | None
|
One-sided target PSD array, mutually exclusive with
|
None
|
target_frequencies
|
ndarray | None
|
Frequencies of |
None
|
ar_order
|
int
|
Autoregressive order |
DEFAULT_AR_ORDER
|
ma_order
|
int
|
Moving-average order |
DEFAULT_MA_ORDER
|
line_frequencies
|
list[float] | None
|
Explicit line frequencies to place poles at. |
None
|
line_widths
|
list[float] | None
|
Widths matching |
None
|
detect_lines
|
bool
|
Whether to detect prominent lines from the target. |
False
|
max_lines
|
int
|
Maximum number of detected lines. |
DEFAULT_MAX_LINES
|
line_prominence
|
float
|
Minimum peak-to-median ratio for detection. |
DEFAULT_LINE_PROMINENCE
|
detectors
|
list[str] | None
|
Detector names to generate. |
None
|
sampling_frequency
|
float
|
Sampling frequency in hertz. |
4096.0
|
duration
|
float
|
Nominal generation duration in seconds. |
4.0
|
seed
|
int | None
|
Base seed for the per-detector generators. |
None
|
low_frequency_cutoff
|
float
|
Lower edge of the target band in hertz. |
2.0
|
high_frequency_cutoff
|
float | None
|
Upper edge of the target band in hertz. |
None
|
block_size
|
int
|
Number of samples generated per filter call. |
DEFAULT_BLOCK_SIZE
|
regularization
|
float
|
Relative ridge added to the zero-lag autocovariance. |
DEFAULT_REGULARIZATION
|
continuation_state()
¶
Return the bounded state needed to keep a running stream going.
export_state()
¶
Return a picklable snapshot sufficient to resume the stream.
from_component(component, config)
classmethod
¶
Construct an ARMA simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector ARMA noise with continuity across calls.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield ARMA chunks lazily while preserving filter state.
import_state(state)
¶
Restore a snapshot produced by :meth:export_state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same configuration. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator's settings. |
reset()
¶
Clear filter state and RNGs.
resume_metadata()
¶
Return the metadata a stopped stream must persist to resume.
ARNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate stateful detector noise from an AR model fit to a target PSD.
autoregressive_coefficients
property
¶
Return a copy of the fitted a_1 .. a_p coefficients.
metadata
property
¶
Return metadata describing the fitted AR model.
model_psd_curve
property
¶
Return the fitted model PSD on the band-masked design grid.
resume_metadata_nbytes
property
¶
Serialized size of the resume metadata in bytes.
state_nbytes
property
¶
Serialized size of the continuation state in bytes.
target_psd_curve
property
¶
Return the band-masked target PSD on the design grid.
__init__(*, psd_file, order=DEFAULT_AR_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization=DEFAULT_REGULARIZATION)
¶
Initialize the simulator and fit the AR model once.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_file
|
str | Path
|
Path or bundled preset name of a two-column PSD table. |
required |
order
|
int
|
Autoregressive order |
DEFAULT_AR_ORDER
|
detectors
|
list[str] | None
|
Detector names to generate. |
None
|
sampling_frequency
|
float
|
Sampling frequency in hertz. |
4096.0
|
duration
|
float
|
Nominal generation duration in seconds. |
4.0
|
seed
|
int | None
|
Base seed for the per-detector generators. |
None
|
low_frequency_cutoff
|
float
|
Lower edge of the target band in hertz. |
2.0
|
high_frequency_cutoff
|
float | None
|
Upper edge of the target band in hertz. |
None
|
block_size
|
int
|
Number of samples generated per recursion call. |
DEFAULT_BLOCK_SIZE
|
regularization
|
float
|
Relative ridge added to the zero-lag autocovariance
before the recursion; |
DEFAULT_REGULARIZATION
|
continuation_state()
¶
Return the bounded state needed to keep a running stream going.
The state is the recursion memory of every detector; its size is proportional to the order and independent of the generated span.
export_state()
¶
Return a picklable snapshot sufficient to resume the stream.
from_component(component, config)
classmethod
¶
Construct an AR-noise simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector AR noise with continuity across calls.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield AR-noise chunks lazily while preserving recursion state.
import_state(state)
¶
Restore a snapshot produced by :meth:export_state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same configuration. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator's settings. |
reset()
¶
Clear detector state and RNGs.
resume_metadata()
¶
Return the metadata a stopped stream must persist to resume.
AddLines
¶
Wrap a base simulator and add spectral lines to its output.
metadata
property
¶
Return additive-wrapper metadata.
__init__(base, lines)
¶
Initialize the additive wrapper.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate base noise and add the configured spectral lines.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield additive spectral-line chunks lazily.
reset()
¶
Reset the additive line state and any resettable base state.
BaseNoiseSimulator
¶
Bases: ABC
Abstract base class for noise simulators.
This interface is the stable API through which the upstream gwmock package
interacts with gwmock_noise. Implementations must override :meth:run.
run(config)
abstractmethod
¶
Run the noise simulation with the given configuration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
NoiseConfig
|
Validated noise simulation configuration. |
required |
Returns:
| Type | Description |
|---|---|
SimulationResult
|
Result containing paths to generated outputs and the config used. |
BlipGlitch
dataclass
¶
Bases: GlitchModel
Gaussian-windowed broadband burst.
__post_init__()
¶
Validate blip-specific parameters.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate a Gaussian-windowed white-noise burst.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
ChannelMetadata
dataclass
¶
Typed, machine-checkable description of one channel.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Channel name, unique within one joint realization. |
kind |
ChannelKind
|
Whether the channel is a gravitational-wave strain channel or an auxiliary/witness channel. |
unit |
str
|
Physical unit of the channel's samples (e.g. |
domain |
ChannelDomain
|
Whether the channel's array is a time series or a frequency-domain series. |
sampling_frequency |
float
|
Sampling frequency in hertz. |
dtype |
str
|
The numpy dtype name (e.g. |
ColoredNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate colored detector noise from an input PSD.
metadata
property
¶
Return simulator metadata.
previous_strain
property
¶
Expose continuity buffers for protocol-compatible state inspection.
__init__(*, psd_file=None, psd_schedule=None, psd_array=None, frequencies=None, delta_frequency=None, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, window_duration=DEFAULT_WINDOW_DURATION)
¶
Initialize the simulator.
export_state()
¶
Return a picklable snapshot of the running crossfade stream.
The snapshot persists the bit-generator state and the chunk counter of the stitcher plus the bookkeeping needed to interpolate the PSD schedule on resume. It contains no cached strain: the previous window is regenerated from the bit-generator state.
from_component(component, config)
classmethod
¶
Construct a colored-noise simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector colored noise.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield colored-noise chunks lazily while preserving simulator state.
import_state(state)
¶
Resume the crossfade stream from an :meth:export_state snapshot.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same settings. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator. |
reset()
¶
Clear continuity and RNG state.
CompositeNoiseSimulator
¶
Additively combine multiple component simulators.
metadata
property
¶
Return metadata describing the composed run.
__init__(components, *, detectors, duration, sampling_frequency, seed)
¶
Initialize the composed simulator.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate all component realizations and add them together.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield additive component chunks lazily.
ConfigurableNoiseSimulator
¶
CorrelatedARNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate correlated detector noise with a truncated VMA representation.
metadata
property
¶
Return metadata describing the fitted VMA model.
__init__(*, psd_files, csd_files=None, order=DEFAULT_VMA_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization_epsilon=DEFAULT_REGULARIZATION_EPSILON)
¶
Initialize the simulator and fit truncated VMA taps once.
from_component(component, config)
classmethod
¶
Construct a correlated-AR simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate correlated per-detector noise with stateful continuity.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield correlated AR chunks lazily while preserving innovation state.
reset()
¶
Clear innovation state and RNG.
CorrelatedNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate correlated detector noise from PSD and CSD inputs.
metadata
property
¶
Return simulator metadata.
previous_strain
property
¶
Expose continuity buffers for protocol-compatible state inspection.
__init__(*, psd_files, csd_files=None, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, regularization_epsilon=1e-12, window_duration=DEFAULT_WINDOW_DURATION)
¶
Initialize the simulator.
from_component(component, config)
classmethod
¶
Construct a correlated-noise simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate correlated per-detector noise.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield correlated-noise chunks lazily while preserving simulator state.
reset()
¶
Clear continuity and RNG state.
DefaultNoiseSimulator
¶
Bases: BaseNoiseSimulator
Default noise simulator implementation.
metadata
property
¶
Return metadata describing the current simulator state.
__init__(*, duration=4.0, sampling_frequency=4096.0, detectors=None, seed=None)
¶
Initialize the simulator with protocol-compatible state.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Return Gaussian white-noise strain arrays.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield white-noise strain chunks lazily.
run(config)
¶
Run the noise simulation with the given configuration.
FitError
¶
Bases: ValueError
Raised when a model fit is numerically invalid or unstable.
GlitchDraw
¶
Bases: NamedTuple
One drawn glitch waveform and the parameters that produced it.
What :meth:GlitchModel.draw returns so an injector can record what it
injected, not merely that it injected something. Every field but the
waveform is optional because not every model has the quantity: a parametric
blip with no PSD has no SNR to report, and only the dataset-backed models
draw a Gravity Spy class.
Attributes:
| Name | Type | Description |
|---|---|---|
waveform |
ndarray
|
The glitch strain, sampled at the requested rate. Its first sample lands at the injection time -- the waveform starts there, it does not peak there. |
amplitude |
float | None
|
The amplitude multiplier drawn from the model's amplitude
distribution, or |
glitch_class |
str | None
|
The drawn morphological class (e.g. a Gravity Spy class
name), or |
target_snr |
float | None
|
The optimal SNR the draw was calibrated to before the
amplitude multiplier, or |
realized_snr |
float | None
|
The optimal SNR of the whole of How it relates to |
GlitchModel
dataclass
¶
Base dataclass for transient glitch generators.
detectors scopes the model to a subset of the run's interferometers. It
accepts one name or a list of them, and defaults to None, which is every
interferometer -- what a model without a selector has always meant. Scoping is
what lets one configuration carry a 10 km model for one set of interferometers
and a 15 km model for another, and per-interferometer rates for a network whose
instruments are not equally glitchy (in O3, Fast_Scattering fired about 29 times
more often in L1 than in H1).
A selector is only half of it. :func:validate_glitch_detector_coverage refuses
a run whose models do not between them account for every interferometer, or
whose models color one interferometer against two different PSDs -- see its
docstring for why proceeding was the defect.
__post_init__()
¶
Validate common glitch parameters.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate a single glitch waveform.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
GlitchNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate transient glitches as an additive standalone component.
from_component(component, config)
classmethod
¶
Construct a glitch-only component from one component definition.
GwoscNoiseSimulator
¶
Real-noise simulator that fetches strain data from GWOSC.
Implements the NoiseSimulator protocol so it can be used
interchangeably with synthetic simulators like
ColoredNoiseSimulator and CorrelatedNoiseSimulator.
Fetches real detector strain from the Gravitational-Wave Open Science Centre and applies user-configured filters to exclude segments containing GW signals or data-quality issues.
When cache_dir is set in the config, downloaded HDF5 files
are saved locally and reused — avoiding repeated downloads for
the same GPS interval.
Attributes:
| Name | Type | Description |
|---|---|---|
duration |
float
|
Duration of the configured GPS interval (seconds). |
sampling_frequency |
float
|
Sampling frequency in Hz. |
detectors |
list[str]
|
List of detector prefixes. |
seed |
None
|
Always |
config |
The underlying GWOSC configuration. |
detectors
property
¶
Return the list of detector prefixes.
duration
property
¶
Return the total duration of the configured GPS interval.
metadata
property
¶
sampling_frequency
property
¶
Return the sampling frequency in Hz.
seed
property
¶
Return None — real noise has no controllable random seed.
__init__(config)
¶
Initialize the real-noise simulator.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
GwoscNoiseConfig
|
Configuration specifying GPS range, detectors, sample rate, filtering options, and optional cache directory. |
required |
check_availability(detectors=None)
¶
Probe GWOSC for per-detector strain-data availability.
Convenience passthrough to :meth:GwoscNoiseFetcher.check_availability.
Use this as a pre-flight check: a detector may have a fully clean
interval (no vetoes) yet have no published strain data, in which
case :meth:generate would fail with a clear error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str] | None
|
Detectors to probe. Defaults to all configured
detectors when |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, bool]
|
A dictionary mapping each requested detector to |
dict[str, bool]
|
GWOSC has data covering the interval, |
generate(duration, sampling_frequency, detectors, seed=None)
¶
Fetch clean noise and return the longest contiguous segment per detector.
Real GWOSC noise is fragmented by the excision of GW events and
data-quality vetoes, so the clean data generally consists of
several non-contiguous spans. To honour the NoiseSimulator
contract — a single, gap-free array per detector — this returns
the longest contiguous clean segment for each detector.
Concatenating the fragments instead (the historical behaviour)
splices non-adjacent samples together, creating discontinuities
that show up as sharp transients after whitening. Use
:meth:generate_segments to obtain every clean segment
separately when the full clean data is required.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
duration
|
float
|
Requested duration (ignored; the longest available contiguous clean segment determines the output length). |
required |
sampling_frequency
|
float
|
Requested sampling frequency (must match
the configured |
required |
detectors
|
list[str]
|
Requested detector list (must be a subset of the configured detectors). |
required |
seed
|
int | None
|
Ignored for real noise. |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, ndarray]
|
A dictionary mapping each detector to a 1-D numpy array |
dict[str, ndarray]
|
holding its longest contiguous clean segment. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
generate_segments(duration, sampling_frequency, detectors, seed=None)
¶
Fetch clean noise and return per-detector lists of contiguous segments.
Unlike :meth:generate, this preserves the gap structure of the
real data: each clean segment — the spans left after excluding GW
events and data-quality vetoes — is returned as its own array, in
GPS-time order. Adjacent segments are not contiguous in time,
so callers must treat each segment independently; concatenating
them would splice non-adjacent samples together and introduce
discontinuities at the joins (which manifest as sharp transients
after whitening).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
duration
|
float
|
Requested duration (ignored; the GPS interval from the config determines the available data). |
required |
sampling_frequency
|
float
|
Requested sampling frequency (must match
the configured |
required |
detectors
|
list[str]
|
Requested detector list (must be a subset of the configured detectors). |
required |
seed
|
int | None
|
Ignored for real noise. |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, list[ndarray]]
|
A dictionary mapping each detector to a list of 1-D numpy |
dict[str, list[ndarray]]
|
arrays, one per clean segment, ordered by GPS time. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield clean-noise chunks lazily.
Fetches the longest contiguous clean segment once (see
:meth:generate) and yields it in chunks of chunk_duration
seconds. Chunking the longest contiguous span keeps every chunk
free of the discontinuities that splicing fragments would
introduce.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
chunk_duration
|
float
|
Duration of each yielded chunk in seconds. |
required |
sampling_frequency
|
float
|
Requested sampling frequency (must
match the configured |
required |
detectors
|
list[str]
|
Requested detector list (must be a subset of the configured detectors). |
required |
seed
|
int | None
|
Ignored for real noise. |
None
|
Yields:
| Type | Description |
|---|---|
dict[str, ndarray]
|
Per-detector strain arrays for each chunk. |
InjectGlitches
¶
Wrap a base simulator and inject transient glitches additively.
A glitch model with no network coherence runs an independent Poisson process
per detector, so every detector receives its own event times and waveform
realizations, and GlitchModel.rate is the event rate seen by each individual
detector. Every (model, detector) pair draws from its own random generator,
derived from the top-level seed and the detector name via an independent
SeedSequence stream.
A model carrying a :class:~gwmock_noise.glitches.models.NetworkCoherence runs
one Poisson process for the whole network and draws one waveform per
event, offering it to every detector it applies to. GlitchModel.rate is then
the network event rate, not the per-detector one, and each detector sees it
times the coherence's participation probability. Arrival times and waveforms come
from a stream keyed by the model alone; each detector's participation and
amplitude ratio come from a stream keyed by its own name and the event's ordinal.
Either way, a detector's realization is reproducible and independent of which other detectors are present or the order in which they are requested -- which is what lets two runs over different networks be compared on the channels they share.
A glitch whose waveform extends past the end of a chunk has its remainder carried into the next chunk and replayed in event order, so the injected glitch series is identical, sample for sample, to a single generate call of the same total duration (the per-sample addition order is preserved exactly, not merely up to rounding). A tail is dropped where the data it would continue into does not exist: past the last generated chunk, and past a segment whose successor the caller places at a discontinuous epoch.
Every event that lands in the strain is recorded as it fires, not merely
tallied: see :attr:glitch_events for the cumulative truth catalogue and
:attr:segment_glitch_events for the events the most recent generate call
injected. A row is written where a glitch starts, so a waveform straddling
a chunk boundary appears once, in the chunk that holds its first sample, at
its true time -- the carried-over tail adds samples to the next chunk but no
second row. GLITCH_CATALOGUE_COLUMNS documents the columns and
GLITCH_CATALOGUE_TIME_CONVENTION which instant each time column names.
gps_start is the GPS time of the first sample of the next segment, and it
belongs to the caller. It advances by each generated segment's duration, so a
stream's catalogue reads in real GPS time on its own; a caller that knows
better -- a writer placing non-contiguous segments, or one driving the stream
batch by batch -- assigns it before each generate call and that assignment is
honoured, including on a call that also carries a seed. Only reset, the
explicit start-over, rewinds it to the epoch the wrapper was constructed with.
glitch_events
property
¶
Return every glitch injected since the last (re)seed, in time order.
Copies, so a consumer cannot edit the run's own record of itself.
metadata
property
¶
Return additive-wrapper metadata.
segment_glitch_events
property
¶
Return the glitches the most recent generate call injected.
What a per-batch consumer wants: the rows belonging to the chunk it is about to write, rather than the whole stream's catalogue re-filtered by time. Empty before the first generate call, and after a call that fired no events.
__init__(base, glitch_models, gps_start=0.0)
¶
Initialize the additive glitch wrapper.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate base noise and inject the configured glitches.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield glitch-injected chunks lazily.
reset()
¶
Reset the additive wrapper and any resettable base state.
Rewinds the epoch to the one the wrapper was constructed with, which
_initialize_process does not: this is the caller saying "start over", where a
seed passed to generate is the caller saying "this segment, whose epoch I have
just told you, begins a new realization". The two need different answers, and
conflating them is what put every streamed segment on the wrong epoch.
JointCovariance
dataclass
¶
A one-sided cross-spectral density matrix over a set of channels.
The convention matches :func:numpy.fft.rfftfreq: frequencies is the
one-sided, non-negative frequency grid, and matrix[k] is the
Hermitian (n_channels, n_channels) one-sided cross-spectral density at
frequencies[k], in units of channel_unit**2 / Hz. The diagonal
entry matrix[k, i, i] is the one-sided PSD of channel
channel_order[i] at that frequency.
Attributes:
| Name | Type | Description |
|---|---|---|
frequencies |
ndarray
|
One-sided frequency grid in hertz, |
matrix |
ndarray
|
Cross-spectral density matrices, |
channel_order |
list[str]
|
Channel names, in the order |
JointDummyCorrelatedSimulator
¶
Public dummy backend producing simultaneous, correlated strain and witness channels.
Every channel is temporally white with unit per-sample variance and
coupling correlation with every other channel; there is no spectral
shaping and no physical model. The one-sided cross-spectral density is
therefore flat and known in closed form: for per-sample covariance
Sigma at sampling frequency fs, the one-sided PSD/CSD is the
frequency-independent matrix 2 * Sigma / fs (the same
covariance-to-PSD scale convention already used by
:class:~gwmock_noise.simulators.multichannel.MultichannelNoiseSimulator,
inverted).
channel_metadata
property
¶
Return typed metadata for every strain and witness channel.
metadata
property
¶
Return simulator metadata.
__init__(*, detectors, witnesses, sampling_frequency=DEFAULT_SAMPLING_FREQUENCY, duration=DEFAULT_DURATION, seed=None, coupling=DEFAULT_COUPLING, strain_unit='strain', witness_unit='counts')
¶
Initialize the dummy backend.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str]
|
Strain channel names. Fixed for the lifetime of this instance; later calls may reorder but not change this set. |
required |
witnesses
|
list[str]
|
Witness channel names, disjoint from |
required |
sampling_frequency
|
float
|
Sampling frequency in hertz. |
DEFAULT_SAMPLING_FREQUENCY
|
duration
|
float
|
Nominal generation duration in seconds. |
DEFAULT_DURATION
|
seed
|
int | None
|
Seed for the shared innovation generator. |
None
|
coupling
|
float
|
Equicorrelation coefficient applied between every pair
of channels (strain-strain, witness-witness and
strain-witness alike). Must satisfy |
DEFAULT_COUPLING
|
strain_unit
|
str
|
Unit string recorded for strain channels. |
'strain'
|
witness_unit
|
str
|
Unit string recorded for witness channels. |
'counts'
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
covariance()
¶
Return this backend's exact, closed-form one-sided cross-spectral density.
Every channel is temporally white, so the returned matrix is
frequency-independent: 2 * Sigma / sampling_frequency at every
frequency, where Sigma is the per-sample covariance matrix drawn
from at construction. This is an exact analytic value, not a
statistical estimate off a realization.
generate_joint(duration, sampling_frequency, detectors, witnesses, seed=None)
¶
Generate one simultaneous, correlated strain+witness realization.
generate_joint_stream(chunk_duration, sampling_frequency, detectors, witnesses, seed=None)
¶
Yield simultaneous strain+witness chunks lazily, preserving RNG state.
JointRealization
dataclass
¶
One simultaneous draw of strain and witness channels.
Attributes:
| Name | Type | Description |
|---|---|---|
strain |
dict[str, ndarray]
|
Per-detector strain arrays, one entry per requested detector. |
witness |
dict[str, ndarray]
|
Per-channel witness arrays, one entry per requested witness channel. |
channel_metadata |
dict[str, ChannelMetadata]
|
Typed metadata for every channel in |
provenance |
dict[str, Any]
|
Implementation-defined provenance describing how this realization was produced (backend identity, package version, seed, RNG algorithm, and any upstream configuration actually used). |
JointStrainWitnessSimulator
¶
Bases: Protocol
Structural interface for simulators producing joint strain+witness output.
This is an optional extension of the legacy NoiseSimulator contract:
implementing it changes nothing about NoiseSimulator itself, and a
NoiseSimulator-only consumer is unaffected by its existence. A class
may satisfy both protocols at once.
metadata
property
¶
Return simulator metadata (same role as NoiseSimulator.metadata).
covariance()
¶
Return the joint one-sided cross-spectral density over all channels.
generate_joint(duration, sampling_frequency, detectors, witnesses, seed=None)
¶
Generate one simultaneous strain+witness realization.
generate_joint_stream(chunk_duration, sampling_frequency, detectors, witnesses, seed=None)
¶
Yield simultaneous strain+witness chunks lazily.
Continuity contract: consecutive chunks from this iterator must equal
the same realization a caller would obtain from one seeded
:meth:generate_joint call spanning the combined duration -- the
same contract NoiseSimulator.generate_stream makes for
strain-only output.
LogNormalAmplitudeDistribution
dataclass
¶
MultichannelNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate multichannel noise from a tabulated PSD/CSD matrix.
ar_coefficients
property
¶
Return a copy of the fitted coefficient matrices A_1 .. A_p.
design_frequencies
property
¶
Return a copy of the design-grid frequencies in hertz.
innovation_covariance
property
¶
Return a copy of the fitted innovation covariance.
metadata
property
¶
Return metadata describing the fitted matrix spectral factor.
model_spectral_matrices
property
¶
Return a copy of the fitted cross-spectral matrices on the design grid.
model_spectral_matrix_curve
property
¶
Return the band-masked fitted cross-spectral matrices on the design grid.
resume_metadata_nbytes
property
¶
Serialized size of the resume metadata in bytes.
state_nbytes
property
¶
Serialized size of the continuation state in bytes.
target_spectral_matrices
property
¶
Return a copy of the target cross-spectral matrices on the design grid.
The matrices are in the covariance scale the recursion uses
(one_sided_psd * sampling_frequency / 2) and carry the band mask and
edge taper the fit was built from.
target_spectral_matrix_curve
property
¶
Return the band-masked target cross-spectral matrices on the design grid.
__init__(*, psd_files=None, csd_files=None, target_matrices=None, target_frequencies=None, order=DEFAULT_ORDER, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE, regularization_epsilon=DEFAULT_REGULARIZATION_EPSILON)
¶
Initialize the simulator and fit the matrix spectral factor once.
The cross-spectral matrix may be given either as files (psd_files
plus csd_files, absent off-diagonal pairs mean zero coherence) or as
an in-memory target_matrices array on target_frequencies, which
lets an analytic pair bypass the tabulated-curve interpolation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_files
|
dict[str, str | Path] | None
|
Per-detector PSD table paths, keyed by detector. |
None
|
csd_files
|
dict[str, str | Path] | dict[tuple[str, str], str | Path] | None
|
Cross-spectral table paths, keyed by detector pair or by
|
None
|
target_matrices
|
ndarray | None
|
One-sided cross-spectral matrices of shape
|
None
|
target_frequencies
|
ndarray | None
|
Frequencies of |
None
|
order
|
int
|
Order |
DEFAULT_ORDER
|
detectors
|
list[str] | None
|
Detector names to generate. |
None
|
sampling_frequency
|
float
|
Sampling frequency in hertz. |
4096.0
|
duration
|
float
|
Nominal generation duration in seconds. |
4.0
|
seed
|
int | None
|
Seed for the shared innovation generator. |
None
|
low_frequency_cutoff
|
float
|
Lower edge of the fit band in hertz. |
2.0
|
high_frequency_cutoff
|
float | None
|
Upper edge of the fit band in hertz. |
None
|
block_size
|
int
|
Number of samples generated per recursion call. |
DEFAULT_BLOCK_SIZE
|
regularization_epsilon
|
float
|
Relative ridge, as a fraction of the largest in-band diagonal entry, added as a zero-lag (white) floor to the autocovariance when the in-band target is not positive definite. Zero forbids the ridge and makes such a target a fit failure. |
DEFAULT_REGULARIZATION_EPSILON
|
continuation_state()
¶
Return the bounded state needed to keep a running stream going.
The state is the recursion history plus the pending output; its size is proportional to the order and independent of the generated span.
export_state()
¶
Return a picklable snapshot sufficient to resume the stream.
from_component(component, config)
classmethod
¶
Construct a multichannel simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector multichannel noise with continuity across calls.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield multichannel chunks lazily while preserving recursion history.
import_state(state)
¶
Restore a snapshot produced by :meth:export_state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same configuration. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator's settings. |
reset()
¶
Clear the recursion history, pending output and generator.
resume_metadata()
¶
Return the metadata a stopped stream must persist to resume.
OverlapSaveFirSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate coloured detector noise with a minimum-phase overlap-save FIR.
filter_taps
property
¶
Return a copy of the designed colouring filter taps.
metadata
property
¶
Return simulator metadata.
resume_metadata_nbytes
property
¶
Serialized size of the resume metadata in bytes.
state_nbytes
property
¶
Serialized size of the continuation state in bytes.
target_psd
property
¶
Return a copy of the band-masked target PSD the filter was designed from.
__init__(*, psd_file=None, target_psd=None, target_frequencies=None, filter_length=DEFAULT_FILTER_LENGTH, minimum_phase=True, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, block_size=DEFAULT_BLOCK_SIZE)
¶
Initialize the simulator and design the colouring filter once.
The target may be given either as a file (psd_file) or directly as a
one-sided array (target_psd), so an analytic target can bypass the
tabulated-curve interpolation entirely. When target_frequencies is
omitted for an array target it defaults to the one-sided grid matching
its length at the configured sampling frequency.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
psd_file
|
str | Path | None
|
Path to a two-column PSD table. |
None
|
target_psd
|
ndarray | None
|
One-sided target PSD array, mutually exclusive with
|
None
|
target_frequencies
|
ndarray | None
|
Frequencies of |
None
|
filter_length
|
int
|
Truncation control |
DEFAULT_FILTER_LENGTH
|
minimum_phase
|
bool
|
Whether to apply the cepstral minimum-phase factorisation. |
True
|
detectors
|
list[str] | None
|
Detector names to generate. |
None
|
sampling_frequency
|
float
|
Sampling frequency in hertz. |
4096.0
|
duration
|
float
|
Nominal generation duration in seconds. |
4.0
|
seed
|
int | None
|
Base seed for the per-detector generators. |
None
|
low_frequency_cutoff
|
float
|
Lower edge of the target band in hertz. |
2.0
|
high_frequency_cutoff
|
float | None
|
Upper edge of the target band in hertz. |
None
|
block_size
|
int
|
Overlap-save block length in samples. |
DEFAULT_BLOCK_SIZE
|
continuation_state()
¶
Return the bounded state needed to keep a running stream going.
The state is the filter memory of every detector plus the pending output and block bookkeeping. It contains no generated strain and its size is independent of the generated span.
export_state()
¶
Return a picklable snapshot sufficient to resume the stream.
from_component(component, config)
classmethod
¶
Construct an overlap-save FIR simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector coloured noise with continuity across calls.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield coloured-noise chunks lazily while preserving filter memory.
import_state(state)
¶
Restore a snapshot produced by :meth:export_state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A snapshot from another simulator with the same configuration. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snapshot does not match this simulator's settings. |
reset()
¶
Clear the filter memory, pending samples and random-number generators.
resume_metadata()
¶
Return the metadata a stopped stream must persist to resume.
This is the per-detector bit-generator state plus the settings needed to rebuild the filter, and is reported separately from the continuation state.
ScatteredLightGlitch
dataclass
¶
Bases: GlitchModel
Arch-shaped scattered-light transient with a Gaussian envelope.
__post_init__()
¶
Validate scattered-light parameters.
applies_to(detector)
¶
coloring_reference()
¶
Return a stable identity for the PSD this model colors against.
None for a model that does no coloring, which therefore imposes no noise
floor on the interferometers it claims and cannot contradict another model
about one. Read from psd_file where the model has one, so a third-party
model gets the same treatment as the built-in ones without restating the
attribute name.
Returns:
| Type | Description |
|---|---|
str | None
|
The resolved PSD identity, or |
draw(sampling_frequency, rng=None)
¶
Generate one waveform together with the parameters that produced it.
What the injector calls, so every event can be recorded rather than merely tallied.
generate_waveform remains the extension point it has always been: a
model that overrides it -- including a subclass of a built-in model --
governs what is injected, and its draw is reported with the parameters
blank rather than filled in from the superclass's draw, which produced a
different waveform. Reversing that precedence would let an override
change the strain while the catalogue described the code it replaced.
generate_waveform(sampling_frequency, rng=None)
¶
Generate a chirping scattered-light glitch.
resolve()
¶
Pin any external, mutable dependency to an immutable version.
Parametric models are fully specified by their configuration and have
nothing external to resolve, so the base implementation is a no-op that
returns None. Models backed by a downloaded dataset (e.g.
DeepExtractorGlitch) override this to fetch and return the concrete
version their run is pinned to. Exposing it on the base lets a caller
pin every model in a heterogeneous glitch list uniformly.
serialize()
¶
Return metadata-friendly model parameters.
SchumannNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate correlated strain noise from an isotropic Schumann-resonance model.
metadata
property
¶
Return simulator metadata.
previous_strain
property
¶
Expose continuity buffers for protocol-compatible state inspection.
__init__(*, positions, coupling_files, detectors=None, schumann_params=None, sampling_frequency=4096.0, duration=4.0, seed=None, low_frequency_cutoff=2.0, high_frequency_cutoff=None, window_duration=DEFAULT_WINDOW_DURATION)
¶
Initialize the Schumann simulator.
from_component(component, config)
classmethod
¶
Construct a Schumann-noise simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate correlated Schumann strain noise.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield Schumann-noise chunks lazily while preserving simulator state.
reset()
¶
Clear continuity and RNG state.
spectrum(frequencies)
¶
Return the isotropic Schumann magnetic PSD on a frequency grid.
theoretical_coherence(frequency, detector_a, detector_b)
¶
Return the isotropic Schumann coherence approximation.
SchumannParams
dataclass
¶
Physical parameters for the isotropic Schumann-resonance model.
__post_init__()
¶
Validate the resonance parameter vectors.
SimulationResult
dataclass
¶
Result of a noise simulation run.
Attributes:
| Name | Type | Description |
|---|---|---|
output_paths |
dict[str, Path]
|
Paths to generated output files, keyed by detector name. |
config |
NoiseConfig
|
The configuration used for the simulation. |
SpectralLineSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate additive spectral lines directly in the time domain.
metadata
property
¶
Return simulator metadata.
__init__(*, lines, detectors=None, sampling_frequency=4096.0, duration=4.0, seed=None)
¶
Initialize the line generator.
from_component(component, config)
classmethod
¶
Construct a spectral-line simulator from one component definition.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate line-only strain for all requested detectors.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield spectral-line chunks lazily.
reset()
¶
Clear continuity state and phase initialization.
WhiteNoiseSimulator
¶
Bases: ConfigurableNoiseSimulator
Generate Gaussian white noise as a composable component.
metadata
property
¶
Return metadata describing the white-noise component.
__init__(*, duration=4.0, sampling_frequency=4096.0, detectors=None, seed=None)
¶
Initialize the white-noise component state.
from_component(component, config)
classmethod
¶
Construct one white-noise component from generic config.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Return Gaussian white-noise strain arrays.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield white-noise strain chunks lazily.
apply_segment_gps_start(simulator, gps_start)
¶
Point every glitch injector inside a simulator chain at one segment epoch.
A writer that assigns a per-segment epoch -- FrameWriter, whose frame names
carry it -- is the authority on where a segment sits in GPS time, and the glitch
truth catalogue has to be stamped with that same instant or the catalogue and the
frame name describe different data. Threading the writer's epoch through is what
keeps it one number rather than two that agree only while the segments happen to be
contiguous.
Walks the wrapper chain (base) and a composite's components, because the
injector is rarely the outermost simulator: a glitch component sits inside
CompositeNoiseSimulator, which sits inside whatever adapter the writer holds.
A simulator with no injector anywhere beneath it is left alone.
A wrapper that builds its simulators rather than holding one -- ParallelAdapter,
which constructs and caches one per detector from a factory -- cannot be reached by
walking attributes, and an epoch set on such a wrapper alone reached none of its
workers. Those declare a writable segment_gps_start and own the fan-out, so this
hands the epoch over and lets them place both the workers they already hold and any
they build later in the same segment.
available_joint_backend_names()
¶
Return the registered joint backend names, sorted.
discover_joint_backends()
cached
¶
Return the entry points registered under :data:JOINT_BACKEND_ENTRY_POINT_GROUP.
This performs discovery only -- it does not import any backend's module.
Call :meth:importlib.metadata.EntryPoint.load on a returned entry point
(or use :func:load_joint_backend) to import and resolve the backend
class it names.
load_joint_backend(name)
¶
Import and return the backend class registered under name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
The entry-point name a package registered its backend under. |
required |
Returns:
| Type | Description |
|---|---|
Any
|
The backend class (or factory callable) the entry point names. |
Raises:
| Type | Description |
|---|---|
KeyError
|
If no backend is registered under |
open_stream(simulator, *, chunk_duration, sampling_frequency, detectors, seed=None)
¶
Open a stateful chunk stream on a protocol-compatible simulator.
take(stream, total_duration, chunk_duration, sampling_frequency)
¶
Collect stream chunks up to total_duration seconds.
whittle_levinson_factorization(autocovariance, order)
¶
Factorize a matrix autocovariance with the Whittle block recursion.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
autocovariance
|
ndarray
|
The autocovariance matrices |
required |
order
|
int
|
Model order |
required |
Returns:
| Type | Description |
|---|---|
WhittleFactorization
|
The fitted autoregressive coefficients, innovation covariance and the |
WhittleFactorization
|
conditioning diagnostics. |
Raises:
| Type | Description |
|---|---|
FitError
|
If the sequence is too short, not finite, or not positive definite at some order of the recursion. |
ValueError
|
If |
Joint-backend conformance suite¶
The shared contract tests a JointStrainWitnessSimulator backend runs under its
own package's test runner (requires pytest).
A shared conformance suite for :class:JointStrainWitnessSimulator backends.
Every package that registers a joint strain+witness backend can run exactly the same contract checks against it, under its own test runner. gwmock-noise runs this suite against its public dummy backend; a downstream package runs it against its own backends.
Usage, in a consumer's test module::
import pytest
from gwmock_noise import load_joint_backend
from gwmock_noise.testing.joint_conformance import JointBackendCase, JointBackendConformance
def _build(seed):
return load_joint_backend("my_backend")(sampling_frequency=64.0, seed=seed)
@pytest.fixture(params=[JointBackendCase("my_backend", _build, 64.0, "my_backend")], ids=lambda case: case.name)
def joint_backend(request):
return request.param
class TestMyBackendConformance(JointBackendConformance):
pass
The suite checks the realization and metadata shapes, seeded determinism, stream continuity, a Hermitian positive-semidefinite covariance over exactly the realized channels, JSON-native provenance and metadata, refusal of an unknown channel, and that no witness is disguised as a strain channel or carried as an array in metadata.
This module requires pytest.
JointBackendCase
dataclass
¶
One joint backend under test, and how to build it.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Label for the case, used in test ids. |
build |
Callable[[int | None], JointStrainWitnessSimulator]
|
Builds a fresh simulator; called with the construction seed, or
|
sampling_frequency |
float
|
The sampling frequency the built simulator runs at. |
provenance_backend |
str
|
The value |
chunk_samples |
int
|
Samples per generated chunk. |
JointBackendConformance
¶
The contract every joint backend must meet.
Subclass it under a Test-prefixed name and provide a joint_backend
fixture yielding :class:JointBackendCase instances.
Beyond the structural protocol, the suite holds a backend to these
conventions: provenance records backend, package_version,
seed and rng_bit_generator; metadata lists the channel names
under detectors and witnesses; and a channel name the simulator was
not built with raises a :class:ValueError whose message says the change
is unsupported.
test_an_unknown_channel_is_refused(joint_backend)
¶
An unknown detector or witness name is refused rather than silently generated.
test_channel_metadata_is_typed_and_complete(joint_backend)
¶
Every realized channel has typed metadata with the right kind, domain, rate and dtype.
test_covariance_is_hermitian_psd_over_the_realized_channels(joint_backend)
¶
The covariance indexes exactly the realized channels and is Hermitian PSD on an rfft grid.
test_no_channel_data_in_metadata_or_provenance(joint_backend)
¶
Metadata, provenance and channel metadata hold no arrays and no channel-length sequences.
test_provenance_and_metadata_are_json_native(joint_backend)
¶
Provenance and metadata survive a strict JSON round trip and record backend, version, seed and RNG.
test_realization_shape_dtype_and_channel_sets(joint_backend)
¶
Strain and witness dicts hold exactly the requested channels as finite 1-D float64 arrays.
test_satisfies_the_public_protocol(joint_backend)
¶
The runtime-checkable public protocol recognizes the backend.
test_seeded_generation_is_deterministic(joint_backend)
¶
Equal seeds give bit-identical realizations; a different seed changes the strain.
test_stream_chunks_continue_the_batch(joint_backend)
¶
Four seeded stream chunks concatenate to the seeded batch over the same span.
test_witnesses_are_never_strain_channels(joint_backend)
¶
Detector and witness names are disjoint, and every witness is typed and delivered as a witness.
array_leaks(value, path='')
¶
Return the paths in a metadata/provenance tree that hold channel-like data.
Flags any :class:numpy.ndarray, and any list or tuple of more than
:data:MAX_METADATA_VECTOR numbers. Sequences of names (channel orders)
are not data and are walked, not flagged.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
Any
|
The tree to walk. |
required |
path
|
str
|
The path of |
''
|
Returns:
| Type | Description |
|---|---|
list[str]
|
One description per leaking path; empty when nothing leaks. |
Parallel execution¶
ParallelAdapter runs independent-detector simulators across threads or
processes (with constraints for correlated backends).
Parallel adapter for independent-detector noise simulators.
ParallelAdapter
¶
Wrap an independent-detector simulator factory with parallel execution.
metadata
property
¶
Return metadata describing the wrapped simulator and executor.
segment_gps_start
property
writable
¶
Return the GPS epoch assigned for the next segment, or None if never set.
__init__(base_factory, *, max_workers=None, backend='auto')
¶
Store the simulator factory and executor preferences.
generate(duration, sampling_frequency, detectors, seed=None)
¶
Generate per-detector strain arrays in parallel.
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield parallel-generated chunks lazily.
reset()
¶
Clear cached worker simulators, resetting continuity across workers.
Diagnostics¶
PSD estimation/comparison and simple stationarity/gaussianity checks used in tests and notebooks.
Diagnostics helpers for gwmock-noise.
DiagnosticResult
dataclass
¶
Outcome of a statistical diagnostic check.
compare_psd(data, target_psd_file, sampling_frequency, rtol=0.1, fmin=10.0, fmax=None)
¶
Compare an estimated PSD against a target PSD file over a frequency band.
This helper compares the median PSD level in-band rather than every bin individually, which makes it robust to the residual variance of a single Welch estimate from stochastic noise.
estimate_psd(data, sampling_frequency, segment_duration=4.0, overlap=0.5)
¶
Estimate a one-sided PSD in physical density units.
The returned frequencies are in Hz. For strain input data, the PSD is in strain^2/Hz.
plot_psd(frequencies, psd, ax=None)
¶
Plot a PSD on log-log axes, importing matplotlib only when needed.
run_diagnostics(data, sampling_frequency, whiten=True)
¶
Run the default Gaussianity and stationarity diagnostics.
Data is whitened against its own Welch PSD estimate before testing (see
_whiten_samples), which makes both checks applicable to coloured
noise. Pass whiten=False to test the raw samples instead — only
meaningful for broadband data.
Output adapters¶
Lazy exports for GWpyAdapter and FrameWriter (optional gwpy / frame
extras).
Output adapters for gwmock-noise.
__getattr__(name)
¶
Lazily resolve optional adapter exports.
GWOSC real-noise fetching¶
Models, segment filtering, data fetching, and a NoiseSimulator wrapper for
retrieving real detector strain data from GWOSC.
GWOSC real-noise fetching subpackage.
Provides configuration models, segment filtering, and data fetching for retrieving real gravitational-wave detector strain data from the Gravitational-Wave Open Science Centre (GWOSC).
GwoscFilterConfig
¶
Bases: BaseModel
Configuration for filtering GWOSC strain segments.
Attributes:
| Name | Type | Description |
|---|---|---|
filter_types |
list[FilterType]
|
Which categories of segments to exclude. |
far_threshold |
float
|
Maximum false-alarm rate (events/year) for GW event filtering. Events with FAR above this threshold are excluded from filtering (treated as noise). Default 1.0 (one per year). |
event_padding |
float
|
Padding in seconds to apply around each GW event GPS time when building vetosegments. |
dq_flags |
list[str]
|
DQ flag basenames (without detector prefix) to query for
data-quality vetosegments. E.g. |
exclude_hardware_injections |
bool
|
Whether to also exclude segments containing hardware injections. |
GwoscNoiseConfig
¶
Bases: BaseModel
Configuration for fetching real detector noise from GWOSC.
Attributes:
| Name | Type | Description |
|---|---|---|
detectors |
list[str]
|
List of detector prefixes (e.g. |
gps_start |
float
|
GPS start time of the requested data interval. |
gps_end |
float
|
GPS end time of the requested data interval. |
sample_rate |
float
|
Desired sampling rate in Hz. GWOSC typically provides 4096 Hz, with some event datasets at 16384 Hz. |
filters |
GwoscFilterConfig
|
Filter configuration for excluding segments. |
host |
str
|
GWOSC host URL. |
cache_dir |
Path | None
|
Optional directory for file-level caching. When set,
downloaded HDF5 frame files are saved locally and reused on
subsequent requests for the same GPS interval. When |
GwoscNoiseFetcher
¶
Fetch real detector noise data from GWOSC with optional filtering.
Downloads strain data from GWOSC and applies user-configured filters to exclude segments containing GW signals and data-quality issues.
When cache_dir is configured, HDF5 files are saved locally and
reused on subsequent requests — avoiding repeated downloads for the
same GPS interval.
Attributes:
| Name | Type | Description |
|---|---|---|
config |
The GWOSC noise fetching configuration. |
clean_segments
property
¶
__init__(config)
¶
Initialize the fetcher.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
GwoscNoiseConfig
|
Configuration specifying detectors, GPS range, sample rate, filtering options, and optional cache directory. |
required |
check_availability(detectors=None)
¶
Probe GWOSC for per-detector strain-data availability.
Unlike :attr:clean_segments, which only reflects vetoes (GW
events and data-quality flags), this checks whether the strain
data itself has been published and is downloadable for the full
configured GPS interval. A detector can have a fully "clean"
interval yet have no published strain — e.g. the data has not
been released yet for the observing run — in which case a fetch
would fail. Use this for a pre-flight check before fetching.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str] | None
|
Detectors to probe. Defaults to all configured
detectors when |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, bool]
|
A dictionary mapping each requested detector to |
dict[str, bool]
|
GWOSC has data URLs covering the interval, |
fetch_clean(detectors=None)
¶
Fetch clean noise segments.
Clean segments are computed by excluding GW events and data-quality issues according to the filter configuration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str] | None
|
Detectors to fetch. Defaults to all configured
detectors when |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, list[TimeSeries]]
|
A dictionary mapping each requested detector to a list of |
dict[str, list[TimeSeries]]
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If no data is available for any requested detector or no clean segments are found. |
fetch_raw(detectors=None)
¶
Fetch raw strain data without filtering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str] | None
|
Detectors to fetch. Defaults to all configured
detectors when |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, TimeSeries]
|
A dictionary mapping each requested detector to a full-interval |
dict[str, TimeSeries]
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If no data is available for any requested detector. |
GwoscSegmentFilter
¶
Compute clean analysis segments from GWOSC metadata.
Queries GWTC event catalogs and data-quality flags to build vetosegments (time windows to exclude) and returns the remaining clean intervals suitable for noise analysis.
Attributes:
| Name | Type | Description |
|---|---|---|
config |
The filter configuration. |
__init__(config)
¶
Initialize the segment filter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
GwoscFilterConfig
|
Filter configuration specifying which segment categories to exclude and their parameters. |
required |
compute_clean_segments(gps_start, gps_end, detectors)
¶
Compute clean (analysis-ready) segments for each detector.
Clean segments are the requested GPS range minus the union of: - GW event vetosegments (if any GW filter is active) - DQ vetosegments (if the data-quality filter is active) - Hardware injection segments (if enabled)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gps_start
|
float
|
GPS start of the requested interval. |
required |
gps_end
|
float
|
GPS end of the requested interval. |
required |
detectors
|
list[str]
|
List of detector prefixes. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, SegmentList]
|
A dictionary mapping each detector to a list of clean |
dict[str, SegmentList]
|
|
get_dq_vetosegments(gps_start, gps_end, detector)
¶
Query data-quality flags for detector and build vetosegments.
Uses gwosc.timeline.get_segments() to fetch pre-computed
DQ veto segments for each flag in dq_flags.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gps_start
|
float
|
GPS start of the query interval. |
required |
gps_end
|
float
|
GPS end of the query interval. |
required |
detector
|
str
|
Detector prefix (e.g. |
required |
Returns:
| Type | Description |
|---|---|
SegmentList
|
A list of |
SegmentList
|
time windows with data-quality issues. |
get_gw_vetosegments(gps_start, gps_end)
¶
Query GWTC events in the GPS range and build vetosegments.
For HIGH_CONFIDENCE_GW, only events with FAR <= far_threshold
are included. For ALL_GW_SIGNALS, all events in the range are
included.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gps_start
|
float
|
GPS start of the query interval. |
required |
gps_end
|
float
|
GPS end of the query interval. |
required |
Returns:
| Type | Description |
|---|---|
SegmentList
|
A list of |
SegmentList
|
on a GW event GPS time with event_padding on both sides. |
get_hardware_injection_segments(gps_start, gps_end)
¶
Query hardware injection events and build vetosegments.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gps_start
|
float
|
GPS start of the query interval. |
required |
gps_end
|
float
|
GPS end of the query interval. |
required |
Returns:
| Type | Description |
|---|---|
SegmentList
|
A list of |
SegmentList
|
hardware injection times. |
Real-noise simulator backed by GWOSC strain data.
Provides a :class:GwoscNoiseSimulator that satisfies the
NoiseSimulator protocol, allowing GWOSC-fetched noise to be used
interchangeably with the built-in synthetic simulators.
GwoscNoiseSimulator
¶
Real-noise simulator that fetches strain data from GWOSC.
Implements the NoiseSimulator protocol so it can be used
interchangeably with synthetic simulators like
ColoredNoiseSimulator and CorrelatedNoiseSimulator.
Fetches real detector strain from the Gravitational-Wave Open Science Centre and applies user-configured filters to exclude segments containing GW signals or data-quality issues.
When cache_dir is set in the config, downloaded HDF5 files
are saved locally and reused — avoiding repeated downloads for
the same GPS interval.
Attributes:
| Name | Type | Description |
|---|---|---|
duration |
float
|
Duration of the configured GPS interval (seconds). |
sampling_frequency |
float
|
Sampling frequency in Hz. |
detectors |
list[str]
|
List of detector prefixes. |
seed |
None
|
Always |
config |
The underlying GWOSC configuration. |
detectors
property
¶
Return the list of detector prefixes.
duration
property
¶
Return the total duration of the configured GPS interval.
metadata
property
¶
sampling_frequency
property
¶
Return the sampling frequency in Hz.
seed
property
¶
Return None — real noise has no controllable random seed.
__init__(config)
¶
Initialize the real-noise simulator.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
GwoscNoiseConfig
|
Configuration specifying GPS range, detectors, sample rate, filtering options, and optional cache directory. |
required |
check_availability(detectors=None)
¶
Probe GWOSC for per-detector strain-data availability.
Convenience passthrough to :meth:GwoscNoiseFetcher.check_availability.
Use this as a pre-flight check: a detector may have a fully clean
interval (no vetoes) yet have no published strain data, in which
case :meth:generate would fail with a clear error.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detectors
|
list[str] | None
|
Detectors to probe. Defaults to all configured
detectors when |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, bool]
|
A dictionary mapping each requested detector to |
dict[str, bool]
|
GWOSC has data covering the interval, |
generate(duration, sampling_frequency, detectors, seed=None)
¶
Fetch clean noise and return the longest contiguous segment per detector.
Real GWOSC noise is fragmented by the excision of GW events and
data-quality vetoes, so the clean data generally consists of
several non-contiguous spans. To honour the NoiseSimulator
contract — a single, gap-free array per detector — this returns
the longest contiguous clean segment for each detector.
Concatenating the fragments instead (the historical behaviour)
splices non-adjacent samples together, creating discontinuities
that show up as sharp transients after whitening. Use
:meth:generate_segments to obtain every clean segment
separately when the full clean data is required.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
duration
|
float
|
Requested duration (ignored; the longest available contiguous clean segment determines the output length). |
required |
sampling_frequency
|
float
|
Requested sampling frequency (must match
the configured |
required |
detectors
|
list[str]
|
Requested detector list (must be a subset of the configured detectors). |
required |
seed
|
int | None
|
Ignored for real noise. |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, ndarray]
|
A dictionary mapping each detector to a 1-D numpy array |
dict[str, ndarray]
|
holding its longest contiguous clean segment. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
generate_segments(duration, sampling_frequency, detectors, seed=None)
¶
Fetch clean noise and return per-detector lists of contiguous segments.
Unlike :meth:generate, this preserves the gap structure of the
real data: each clean segment — the spans left after excluding GW
events and data-quality vetoes — is returned as its own array, in
GPS-time order. Adjacent segments are not contiguous in time,
so callers must treat each segment independently; concatenating
them would splice non-adjacent samples together and introduce
discontinuities at the joins (which manifest as sharp transients
after whitening).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
duration
|
float
|
Requested duration (ignored; the GPS interval from the config determines the available data). |
required |
sampling_frequency
|
float
|
Requested sampling frequency (must match
the configured |
required |
detectors
|
list[str]
|
Requested detector list (must be a subset of the configured detectors). |
required |
seed
|
int | None
|
Ignored for real noise. |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, list[ndarray]]
|
A dictionary mapping each detector to a list of 1-D numpy |
dict[str, list[ndarray]]
|
arrays, one per clean segment, ordered by GPS time. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
generate_stream(chunk_duration, sampling_frequency, detectors, seed=None)
¶
Yield clean-noise chunks lazily.
Fetches the longest contiguous clean segment once (see
:meth:generate) and yields it in chunks of chunk_duration
seconds. Chunking the longest contiguous span keeps every chunk
free of the discontinuities that splicing fragments would
introduce.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
chunk_duration
|
float
|
Duration of each yielded chunk in seconds. |
required |
sampling_frequency
|
float
|
Requested sampling frequency (must
match the configured |
required |
detectors
|
list[str]
|
Requested detector list (must be a subset of the configured detectors). |
required |
seed
|
int | None
|
Ignored for real noise. |
None
|
Yields:
| Type | Description |
|---|---|
dict[str, ndarray]
|
Per-detector strain arrays for each chunk. |
Command-line interface¶
Typer application and the simulate command used by the gwmock-noise console
script.
Main entry point for the gwmock_noise CLI application.
main(verbose=LoggingLevel.INFO)
¶
Implement the main entry point for the CLI application.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
verbose
|
Annotated[LoggingLevel, Option('--verbose', '-v', help='Set verbosity level.')]
|
Verbosity level for logging. |
INFO
|
register_commands()
¶
Register CLI commands.
setup_logging(level=LoggingLevel.INFO)
¶
Set up logging with Rich handler.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
level
|
LoggingLevel
|
Logging level. |
INFO
|
Command for running noise simulations.
simulate(config_path)
¶
Run noise simulation from a configuration file.
Loads the configuration, validates it, and runs the noise simulator. Output files are written to the directory specified in the config.