Scientific caveats
Grows with each milestone. Seeds:
- A connectome is wiring; it carries no measured dynamics.
- Neurotransmitter identities are classifier predictions (Eckstein et al. 2024 lineage; MaleCNS body-level file gives argmax + confidence only).
- Weights are synapse counts under a detection and confidence threshold that differs per dataset (see
DATA_SOURCES.md). - Simulation outputs are model predictions; the Shiu et al. constants were fitted to FlyWire, so MaleCNS runs are uncalibrated until stated otherwise.
- Bit-exact determinism holds on CPU float64 only; MPS/CUDA results are seeded but summation-order dependent.
Path and cell-type caveats (M2)
- Labellar gustatory receptor neurons: FlyWire v783 types sugar and water GRNs together as
LB3(bitterLB1a/b, Ir94eLB2); MaleCNS splitsLB3a–d. Selecting "sugar GRNs" by type therefore includes water GRNs; use explicit ID lists (e.g. Shiu et al. 2024) where sugar specificity matters. - Path "strength" is a product of per-hop input fractions; totals include synapses from unannotated fragments where the dataset provides them (MaleCNS), so fractions are smaller than in neuron-only graphs and are not comparable across datasets without care.
- The Shiu et al. export carries only a per-neuron sign, not a transmitter; flyconn stores
acetylcholine/gabaplaceholders so that sign policies reproduce the model, and labels the sourceshiu2024_excitatory_flag.
Uncertainty caveats (M3)
- Confidence-only NT model. Where a dataset gives only an argmax label and a confidence (MaleCNS body-level file without the opt-in tbar aggregation), flyconn puts the confidence on the label and spreads the remainder uniformly over the other classes. This is a modelling convenience, not a calibrated posterior; results are labelled with the model used per neuron (
nt_probabilitiesreturns it). - Per-body mean of per-synapse probabilities (MaleCNS
--level nt-probs, FlyWire Codex*_avg) ignores within-neuron heterogeneity and synapse-level confidence; Dale's-law single sign per neuron is assumed throughout. - Degree-preserving rewiring keeps each neuron's out-degree, its outgoing weight multiset and every neuron's in-degree, but not in-weight or spatial structure; it is a configuration-model null, not a biological alternative wiring.
- Version drift by id.
diff_versionsmatches neurons by root/body id only. FlyWire root ids change with proofreading (only 106,785 of 127,978 v630 ids persist in v783), so absent ids are reported as unmatched rather than silently treated as removed neurons. - Empirical p-values use the (k+1)/(n+1) correction and are two-sided around the null mean; with 100 nulls the smallest reportable p is about 0.01.
Simulation caveats (M4)
- The LIF model has one free parameter (w_syn = 0.275 mV) that Shiu et al. chose so that 100 Hz sugar-GRN drive gives about 80 % of MN9's maximal rate in FlyWire v630. Every constant was fitted to FlyWire; MaleCNS (and any other dataset) runs are uncalibrated and their spike counts are not comparable until the re-calibration protocol has been applied and reported.
- Baseline firing is 0 Hz; there is no intrinsic activity, noise, neuromodulation, gap junctions, or synaptic plasticity. Outputs are model predictions of wiring-constrained excitability, not predictions of behaviour.
- Signs come from predicted transmitters under a single-sign-per-neuron rule; histamine is inhibitory by default in flyconn but Shiu's FlyWire predictions had no histamine class.
- Exactness: bit-identical spikes are guaranteed only for the CPU float64 path. float32 on CPU/MPS/CUDA matched the reference in every test so far, but summation order can move borderline spikes by a tick; compare distributions, not single runs.
- Tick kernel: large runs (and any run with
compile=True) fuse the per-tick update withtorch.compile; provenance recordstick_kernel(torch.compileoreager). On CPU, inductor builds without FP contraction or unsafe math, and the compiled kernel reproduces the eager kernel and the v0.2.0 per-tick loop spike for spike (unit tests, plus the full v630 network in float32 and float64). MPS matched on the same full-network check but is not guaranteed; on CUDA, Triton may fuse multiply-adds, so compiled float32 can differ from eager at rounding level (the same class of difference as summation order). If compilation fails, the engine falls back to eager with a warning;compile=Falseforces eager.
Experiment and report caveats (M5)
- Readout statistics compare per-trial mean rates of a readout group between conditions. Trials share the connectome and differ only in Poisson input seeds, so "n" is the number of trials, not of animals; effect sizes describe the model, not biology.
- Controls are simulated with the same seeds on degree-preserving rewired and sign-shuffled connectomes; a stimulated-vs-control difference shows that the specific wiring (or the specific signs) matters for the readout under this model, nothing more.
- Benjamini-Hochberg correction is applied across all readout x comparison rows of one experiment; with few rows the correction is mild.
- Sign shuffling as a null: with roughly two thirds of neurons excitatory, permuting signs across neurons breaks excitation/inhibition balance and typically produces runaway activity (seen in the W2 run: 59M spikes vs 0.6M). Treat it as a test of whether inhibitory identities matter where they are, not as a matched-activity null; compare readouts under the rewired control first.
- Perturbations are all-or-none (silencing zeroes outgoing synapses; stimulation is Poisson kicks); they do not model partial knockdown, developmental compensation or neuromodulation.
Sweep caveats
- A sweep (
sweep: {silence_each: ...}orsweep: {rate_hz: [...]}) runs the shared conditions (baseline, stimulated, controls) once and one simulation per variant. Every variant reuses the spec's seeds, so variant and reference trials share Poisson input; the Mann-Whitney test treats them as unpaired, which is conservative. - n is the number of trials per variant. Benjamini-Hochberg runs across all variant x readout rows of the sweep (
sweep.parquet); the shared stimulated-vs-baseline/control comparisons inreadouts.parquetform a separate family. silence_eachsimulates only themax_itemscandidate groups with the highest summed firing rate in the stimulated condition (in the spirit of Shiu et al. 2024's screen of the most active neurons). Groups that never fire are listed as skipped: silencing zeroes outgoing synapses, so they cannot change that run. A type that is silent under this stimulation may still matter under another one. Ranking by summed rate favours types with many neurons.- A screen ranks effects within one model and one stimulation; it is a list of candidates for experiments, not evidence that a type is necessary in the fly. With an uncalibrated dataset (MaleCNS) the magnitudes are not comparable to FlyWire runs.
- To bound file size,
counts.parquetstores all non-zero counts for the shared runs but only the readout neurons' counts for variant runs; per-variant totals are inconditions.parquetandvariants.parquet. - In a
rate_hzsweep everystimulateentry is set to the same rate; the reference is the unstimulated baseline (or the stimulated condition when the baseline is disabled).
Cross-dataset comparison caveats (M6)
- The null model is left/right variability within a brain. Schlegel et al. 2024 report that connectivity between two different female brains (FlyWire vs hemibrain) is less similar than left vs right within one brain (cosine effect size 0.045 ± 0.096). A male-vs-female difference must therefore be judged against that between-individual range, not only against the L/R permutation p-value;
compare_typegrades its verdict accordingly (no evidence / detectable but within the reported between-brain range / beyond it). - Partner types without a cross-dataset match (e.g. VNC neurons, which FlyWire's brain volume cannot contain, or unmatched types) are excluded from the profiles and their fraction is reported; a large unmatched fraction means the comparison covers only part of the type's output or input.
- Datasets differ in synapse detection, confidence thresholds and proofreading completeness; use the same
min_weighton both sides and expect residual technical differences. - Sex-specific types by definition have no counterpart;
match_typescannot represent them, and dimorphism annotations should be consulted separately (MaleCNSdimorphism, FlyWiredimorphism).
Genetic access caveats (M7, exploratory)
- Line-to-type annotations (Meissner et al. 2025) are free text across nomenclatures (hemibrain, MANC, optic lobe, FlyWire) and carry a 1-4 quality grade; expression confidence lives in each line's source paper. EM body ids in the table have no dataset tag.
- NeuronBridge hits are colour-depth similarity candidates ranked by score; a hit is not an expression call, scores are not comparable across libraries, and absence of a hit is weak evidence. Off-target lists are long by construction; rank and inspect.
- Most FlyLight imagery is female and adult with specific reporters; expression can differ in males or with other effectors.
BANC caveats (M8)
- The released
banc_888_edgelist_simple_v2.feather(checksum-pinned, updated 2026-06-02) has 11,752,828 rows, not the 11,510,975 its documentation states, and contains 156,311 self-connections although the documentation says autapses were removed. flyconn drops them and reports the count. - BANC's
malecns_cell_typecolumn refers to MaleCNS v0.9; flyconn stores it asmalecns_09_cell_typeand does not treat it as a v1.0 mapping. - Comparing MaleCNS with BANC (both brain plus nerve cord) is the more complete male-vs-female test than MaleCNS vs FlyWire (brain only). For DNp01 the unmatched output fraction drops from 58 % (vs FlyWire) to 0 % (vs BANC). Two female datasets can also disagree: EPG was graded "beyond the between-brain range" against FlyWire (statistic 0.189) but only "within range" against BANC (0.088); when two female references disagree, the difference is not attributable to sex.
- Both antennal nerves were damaged in the BANC sample (Bates et al. 2026); antennal inputs are under-represented.
Hemibrain and MANC caveats
- Hemibrain weights are neuPrint
weight, which counts every postsynaptic density because the dataset'spostHighAccuracyThresholdis 0; the high-confidenceweightHPis not in the export. The store is the v1.2 export; neuPrint v1.2.1 differs by 342 edges / 4,077 synapses. - The hemibrain volume is mostly the right hemisphere. Left-side neurons and left-side arbors are truncated, so the left/right null used by
compare_typecan be meaningless for a hemibrain type (PFL3: within-hemibrain left/right similarity 0.31 against 0.99 in MaleCNS). Checkwithin_similarity_bbefore reading a verdict. - Hemibrain side, ItoLee hemilineage and short
cell_classcome from Schlegel et al. 2024 Supplemental file 5 (411 traced bodies have no row);super_class,flowandregioncome from the sjcabshemibrain_121_meta. No FlyWire type cross-reference is filled because none of these files publishes one, so hemibrain partners keep hemibrain names under every partner vocabulary. - Hemibrain transmitter probabilities come from
hemibrain-v1.2-body-mean-neurotransmitters.feather, uploaded without a README; its provenance is UNVERIFIED (the class set matches the Eckstein et al. 2024 classifier). All 1,927 typed Kenyon cells (KC*) are predicted dopaminergic, although Kenyon cells are cholinergic; Eckstein et al. 2024 report the same misprediction. Prefer per-type known transmitters for KCs. - MANC has no flat v1.2.x export. flyconn uses the sjcabs compiled meta and edge list (built by flyconnectome/bancpipeline); their row count and synapse sum differ slightly from neuPrint, and the confidence threshold used for the build is not documented. Per-neuron input totals count traced partners only.
- MANC transmitter predictions are 3-class (acetylcholine, GABA, glutamate); a MANC neuron is never predicted monoaminergic.
nt_confis the Codex confidence and is filled only where Codex's label agrees with the sjcabs label (23,410 of 23,650). - MANC cell types are neuPrint
typenames (e.g.Ti flexor MN); Codex's primary types differ for 420 neurons.manc_121_cell_typeequalscell_type, so MANC partners share the FlyWire-or-MANC vocabulary with MaleCNS and BANC.