Input and output#
PyHermes has two I/O layers. Particle readers turn catalogue files into a
common {"pos": (N, 3), ...} structure. Task-specific data classes read and
write reusable MRA fields and statistical products.
Particle input#
SFCProjection.fin selects a reader:
fin:
path: "./data/catalog.bin"
format: "bin"
reader_params: {}
catalog_weight_key: null
field_value_key: null
catalog_weight_key selects completeness or selection weights.
field_value_key selects a per-particle mark or physical quantity. Both must
resolve to one-dimensional arrays with one entry per particle. null means
unit values.
Remote files and caching#
fin.path may be a local path or an HTTP(S) URL. Remote inputs are resolved
to a local file before the format-specific reader runs:
fin:
path: "https://data.example.org/haloes.npz"
format: "npz"
download:
cache_path: "./data/haloes.npz"
sha256: "0123456789abcdef..."
timeout: 60
cache_path gives one exact destination. Alternatively, cache_dir puts
the file in that directory under a deterministic URL-derived name. With
neither option, PyHermes uses $PYHERMES_CACHE_DIR, then
$XDG_CACHE_HOME/pyhermes, and finally ~/.cache/pyhermes.
An existing cache file is reused after optional SHA256 verification. New data
are downloaded to a unique partial file and atomically renamed only after a
successful transfer and checksum. Under MPI, SFCProjection performs this
work only on rank 0 before distributing particle slabs. Remote input targets
single files such as NPZ, BIN, or a complete group_tab file. Directory
catalogues and split datasets should remain on local or shared storage.
The Particle I/O notebook starts from the
public Quijote FoF file, constructs local NPZ and BIN equivalents, and
displays their array names, shapes, dtypes, and units beside the conversion
code. These converted files demonstrate supported catalogue formats; the FoF
file is the only distributed source catalogue. The data stay outside Git,
while
examples/configs/param_sfc_projection.yaml records the stable URL and
checksum used by the Quick Start.
Supported formats#
|
Reader |
|---|---|
|
Raw binary table with configurable dtype and column selectors. |
|
NumPy archive with a position key and optional field-key mapping. |
|
Legacy Gadget-format snapshot, including split files. |
|
Single or split Gadget HDF5 snapshot with explicit unit scales. |
|
Legacy Gadget FoF catalogue without SUBFIND. |
|
Quijote/Pylians-style |
A recognized local or URL-path suffix can infer bin, npz, or
gadget when format is omitted. The group_tab_###.# filename
pattern similarly infers fof. Directory and HDF5 base paths should set the
format explicitly.
Raw binary tables#
fin:
path: "./data/haloes.bin"
format: "bin"
reader_params:
dtype: "float32"
ncols: 7
pos_cols: [0, 1, 2]
fields:
vx: 3
vy: 4
vz: 5
mass: 6
catalog_weight_key: null
field_value_key: "mass"
A scalar selector returns one column; a list selector returns multiple columns. Only scalar fields can be selected as catalogue weights or projected field values.
NPZ archives#
fin:
path: "./data/haloes.npz"
format: "npz"
reader_params:
pos_key: "position"
fields:
completeness: "weight"
vx: "velocity_x"
catalog_weight_key: "completeness"
field_value_key: "vx"
When fields is omitted, every array except pos_key is exposed under its
original name.
Gadget snapshots#
The HDF5 reader resolves either one file or a base path such as snap_004
with siblings snap_004.0.hdf5, snap_004.1.hdf5, and so on.
fin:
path: "/path/to/snapdir_004/snap_004"
format: "gadget_hdf5"
reader_params:
ptype: 1
position_scale: 1.0e-3
velocity_scale: 1.0
mass_scale: 1.0
load_velocity: false
load_mass: false
For Quijote coordinates stored in \(\mathrm{kpc}/h\),
position_scale=1e-3 converts to \(\mathrm{Mpc}/h\). Velocity and mass
are not loaded unless requested, which avoids large unused arrays. h5py is
required; compressed snapshots may also require hdf5plugin.
Quijote FoF catalogues#
fin:
path: "https://pyhermes.astroslacker.com/downloads/group_tab_004.0"
format: "fof"
download:
cache_path: "./data/quijote_halos/8000/groups_004/group_tab_004.0"
sha256: "4a1c6ca4f6747a70e9e552685226ecf5d678c6c97551e2caa7cc3883502eac85"
reader_params:
redshift: 0.0
fields:
mass: "mass"
vx: "vel_x"
vy: "vel_y"
vz: "vel_z"
catalog_weight_key: null
field_value_key: null
The reader exposes pos, mass, vel, component velocities, npart,
and group_offset before optional selection. It converts Quijote positions
to \(\mathrm{Mpc}/h\), masses to \(M_\odot/h\), and applies the reader’s
redshift velocity convention. A direct file whose header declares multiple
pieces discovers .1, .2, and later siblings beside the supplied
.0 file. Alternatively, pass the catalogue root as path and set
reader_params.snapnum; this preserves the existing
groups_004/group_tab_004.0 directory convention.
In-memory input#
File input can be replaced by arrays:
task = SFCProjection()
task.particle_pos = pos
task.catalog_weight = completeness
task.field_value = mass
task.weight_normalization = "raw"
field = task.run(save_result=False)
Positions must have shape (N, 3). Scalar weights or field values are
broadcast; arrays must be finite and have shape (N,).
Saved data classes#
Task |
Reader class |
Main arrays |
|---|---|---|
|
|
|
|
|
sampled positions and |
|
|
sampling coordinates, count products, and |
|
|
angles, side geometry, triplet products, |
|
|
samples, radial windows, |
Read products through those classes:
from pyhermes.io import SFCField, Corr2PCFData, Corr3PCFMultipoleData
field = SFCField(data_path="./output/halo_sfc.pkl", threads=8)
corr = Corr2PCFData(data_path="./output/halo_2pcf.pkl")
multipoles = Corr3PCFMultipoleData(
data_path="./output/halo_3pcf_multipoles.pkl"
)
The current files contain NumPy-framed, pickle-serialized Python metadata. They are trusted local products, not a safe interchange format for untrusted files. Treat the data-class interface as public and the internal dictionary layout as an implementation detail.
PyHermes converts path-like metadata to strings when writing new products.
The loader also translates the private pathlib._local path classes emitted
by Python 3.13, so trusted products can move between Python 3.12 and 3.13
without requiring users to rewrite task.fin. This targeted compatibility
does not make arbitrary pickle objects portable or safe to load.
Particle companions#
save_particle_data: true writes positions, catalogue weights, and field
values to particle_data_path as NPZ. When the path is empty, it is derived
from fout_path. This companion is required when a later particle-centred
task cannot recover the original catalogue from fin. Derived fields made by
arithmetic or convolution do not preserve a unique particle catalogue; pass
explicit centre arrays when using them as a particle-centred first vertex.
Output policy#
Only rank 0 writes task products. Parent directories are created as needed.
Existing files are protected unless overwrite=True is passed to run.
Store large generated products outside the tracked examples/ source files;
the repository’s scripts and notebooks use examples/output/ for this
purpose.