arXiv Science⌕ Search

arXiv · 2610.09114

A Connectome Test of the Fly Hashing Algorithm

Abstract

Dasgupta, Stevens and Navlakha (2017) showed that the Drosophila olfactory circuit, modelled as a random sparse projection followed by winner-take-all, is a locality-sensitive hash that beats classical LSH. The projection was random because the wiring was unknown. We test it against four electron-microscopy connectomes (MaleCNS, hemibrain, FlyWire, BANC; four animals, seven hemispheres). First, the 2017 pattern holds in a reimplementation of its protocol on SIFT, MNIST and odour mixtures (GloVe is near chance at short codes for every method): the fly hash beats k Gaussian projections at short hash lengths (3.1x in AP@200 on MNIST at k = 4). Second, against the tested real-valued Gaussian baseline that advantage is per active cell, not per operation: Gaussian projections given the same projection arithmetic retrieve better on every dataset and input dimension tested. Third, across four connectomes the measured pairing of glomeruli gives no consistent retrieval advantage over degree-preserving rewiring: retrieval is slightly lower (median -1.6%), and the small odour deficits depend on how missing odour responses are treated. Separately, equalising glomerular fan-out at fixed connection count improves retrieval in the model in every hemisphere, while equalising inputs per cell lowers it on average. Yet the fan-out profile is similar across the four sampled animals (median between-animal Spearman rho = 0.86), structural synapse counts do not offset it, and its relation to odour tuning is weak. In the model that uneven allocation costs retrieval. For practice: measured wiring gives no consistent retrieval advantage over degree-preserving random wiring, so a fly hash needs no connectome data, and its advantage is per active unit, which may suit hardware where active units rather than arithmetic are the binding cost, a hypothesis we do not test.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sebastian Senge. 2026-10-06. A Connectome Test of the Fly Hashing Algorithm. https://arxiv.org/abs/2610.09114

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Relocating Nonlinearity: How a Downstream Learner Reshapes What Genetic Programming Must Evolve

Genetic programming was conceived as a way of evolving solutions: the program is the answer, and fitness is the error of its own output. A substantial line of work instead makes the program an input to a separate learner, so fitness measures the learner's output rather than the program's. Which learner to attach matters, with no account of what decides it. What makes a target hard for genetic programming is how much nonlinearity the program must build by composing primitives; whatever nonlinearity the learner supplies, the program need not. The quantity to measure is therefore how much of a target's nonlinearity a learner can take over, though on continuous benchmarks it can only be estimated. Here we show that how much the program must still build is decided by which learner is attached, and is measurable on the programs themselves. Moving to Boolean domains, where a target's nonlinearity is exactly its Fourier degree, we tune that degree from one to six with everything else fixed. Conditioned on success, a linear learner forces the evolved program to the target's degree exactly, at one through six without exception over thirty runs per setting, while tree ensembles succeed with far simpler programs and four times as often. This gives the field two things: a learner can be chosen from a target's structure instead of its reputation for difficulty, and the degree of the evolved program is a diagnostic free to compute during any run. Our control target has lower degree than the parity problems yet gains nothing from a nonlinear learner, while targets reducible to a simple statistic gain a great deal. The latter comes with a warning: evolution internalises what the learner supplies in only four of sixty-three conditions, and grows more dependent on it in forty-four, so the more capable the learner, the less of the model is legible in the program.

cs.NE↗

The Modular CMA-ES: A Framework for Modern Evolution Strategies

Since their introduction, modern evolution strategies, such as the Covariance Matrix Adaptation Evolution Strategy (CMA-ES), have become established as powerful methods for continuous black-box optimization. This success has led to a wide range of proposed modifications, each designed to improve performance or behavior in specific optimization scenarios. However, because these developments have largely been introduced and studied in isolation, their interactions remain comparatively underexplored. In this paper, we present the Modular CMA-ES (ModCMA), a configurable framework that integrates a wide range of mechanisms from modern evolution strategies within a single implementation. By decomposing CMA-ES into modules with interchangeable options for sampling, selection and recombination, step-size adaptation, matrix adaptation, and restarting, ModCMA enables systematic exploration of a large design space of modern evolution strategies and facilitates the construction, comparison, and automated configuration of new algorithm variants. We illustrate the benefits of this modular approach through two example studies. First, we compare several matrix-adaptation mechanisms in terms of their computational cost and optimization performance. Second, we use automated algorithm configuration to specialize ModCMA to individual benchmark problems and analyze the resulting configurations. Together, these examples demonstrate how the framework can be used both to study individual algorithmic design choices and to explore their combinations in a systematic and reproducible manner.

cs.NE↗

Evolutionary Architecture Search for Chlorophyll-$a$ Prediction in Lakes using Sentinel-2

Small tabular datasets with expert-designed spectral features are the norm in operational Earth observation, and the networks applied to them are typically hand-designed. We revisit one such published model -- a Sentinel-2 algal bloom classifier -- and ask what architecture search adds, holding the task, the features and the lake-level train/test split of the original study fixed. Searching an extended multilayer-perceptron space with regularized evolution, and selecting on inner-cross-validation AUC only, we find networks that improve held-out AUC from 0.790 to 0.820 and accuracy from 0.733 to 0.748 while using 409 trainable parameters, 26 times fewer than the strongest hand-designed reference. The search converges on a consistent recipe -- a single narrow layer, RMS normalisation, $\tanh$ activation, step-decayed RMSprop and weight averaging -- that a practitioner would be unlikely to reach by default. At 1.6\,kB the resulting model is small enough to serve as an onboard screening trigger, which is the setting that motivates the work. Code: https://github.com/VU-AIML/automl4eo-bloom-nas.

cs.NE↗