Automated Daily Update | Last Run: 2026-08-02 09:54 UTC
Papers are automatically categorized by topic and sorted by date.
2026-07-29 | Artem Litvinenko, Erbin Qiu, Sambit Ghosh et al.
VO$_2$-based relaxation oscillators form a rapidly developing field that finds applications in neuromorphic computing, Ising machines, and numerous signal processing concepts. These oscillators operate in a deeply nonlinear relaxation regime based on rapid phase transitions between insulating and metallic states in the VO$_2$ material. This process is governed by thermal effects, which lead to additional voltage fluctuations and contribute to a considerably wide spectral linewidth in the VO$_2$-based oscillator signal. In this work, we thoroughly study the phase noise in VO$_2$-based relaxation oscillators and demonstrate that the broadening of the generation spectrum linewidth at low oscillation frequencies is caused by an increased susceptibility to thermal fluctuations during the incubation phase. We explore the types of noise affecting oscillator stability and show that synchronization with an external square-wave signal improves the phase noise more effectively than a sinusoidal-shape injection locking signal.
Investigating reservoir computing for branch predictionin pipelined processors using emerging CMOS memristor devices
2026-07-29 | Harvey Samuel George Johnson, Sendy Phang
This project aimed to develop a novel reservoir compute (RC) implementation framework targeting high-speed operation and integration with CMOS digital logic. With the target workload of branch prediction (BP) for multistage pipelined central pro-cessing unit (CPU) cores. For this, a novel memristor based RC design framework was developed within the context of the workload requirements. This was then implemented in simulation using industry standard modelling languages of System Verilog (SV) and Verilog-AMS (VAMS).The developed RC design framework was subsequently verified using a basic sequence detection task before further benchmarking for its effectiveness at BP. The developed RC framework was tested using the Dhrystone performance benchmark, while targeting the RISC-V RV64GC instruction set architecture (ISA). Conducted testing demonstrates that RC shows great promise for ap-plication to BP and is capable of achieving impressive overall prediction accuracy. However, testing also shows that further refinement of the developed RC design framework is necessary to address shortfalls in the adaptability of the proposed RC system. As comparison against the state of the art TAGE predictor showed the proposed RC design framework to be 15x slower to adapt to changes in branching behaviour.
2026-07-29 | Menelaos Skontranis, Benoit Charbonnier, Olivier Girard et al.
In this work we experimentally investigate the spiking dynamics of two-section InP quantum-well lasers monolithically integrated on silicon. By appropriately tuning the electrical bias conditions, we realize multiple neuronal-like operating regimes, such as integrate-and-fire and resonate-and-fire, highlighting the device's versatility as a high-speed photonic neuron. A systematic investigation of laser design parameters, including cavity length and gain/saturable absorber ratio, elucidates their impact on spiking-related properties (such as pulse repetition frequency) and traces the operational parameter space that unlocks stable spiking. Finally, these findings pave the way toward scalable neuromorphic photonic integrated circuits, where low-loss silicon synapses coexist with versatile laser neurons.
2026-07-29 | Lukas Endres, Hannes Töpfer, Michaela Blum et al.
Memristors are promising devices for applications such as non-volatile memory, neuromorphic computing, logic circuits, and analog signal processing. The development of such systems requires accurate simulations based on models that reproduce the electrical behavior of real devices under both continuous and pulsed excitation. This work presents the development of a simulation environment for a volatile TiO2-based memristor. Experimental measurement data are analyzed to verify the memristive behavior of the device and to identify a suitable model. The model parameters are then optimized to match the measured characteristics. The resulting model is implemented in SPICE and validated by comparing simulation results with measurement data. The comparison shows a good agreement between simulation and experiment, demonstrating that the developed model is suitable for reproducing the electrical behavior of the investigated memristor and can be applied in circuit-level simulations, as demonstrated by a leaky integrate-and-fire neuron.
Mitigating the Impact of Retention Loss on Inference Accuracy in 65 nm Single-Poly Floating-Gate Analog In-Memory Computing
2026-07-27 | Mirko Brazzini, Giulio Filippeschi, Alessandro Catania et al.
We show with experiments and system-level simulations that it is possible to successfully mitigate the impact of retention loss on inference accuracy degradation by using both circuit-level compensation techniques and batch normalization recalibration at the algorithmic level. Experiments are performed on a single-poly floating-gate (FG) analog non-volatile memory array for analog in-memory computing fabricated in a standard 65 nm CMOS. We use a model of retention-loss statistics calibrated with experiments to evaluate the system-level impact on neural network models such as VGG-10/CIFAR-10 and WideResNet-28-10/CIFAR-100. We show that, after 60 days since programming, combined mitigation techniques enable to recover the baseline inference accuracy within 2-4%
2026-07-27 | Xiaoyi Lei, Fanfu Wu, Yunting Liu
Embedded fault diagnosis in three-phase inverters must satisfy the sub-watt power budget of converter control hardware, but conventional convolutional neural network (CNN)-based methods require dense multiply-accumulate operations and impose substantial inference energy. This work proposes an event-driven neuromorphic framework for energy-efficient open-circuit (OC) fault diagnosis. A CNN trained on current-vector trajectory matrices is converted into a spiking neural network (SNN) and evaluated using the NengoLoihi framework with Loihi-based neuromorphic energy estimation. By exploiting the sparse structure of trajectory matrices, the SNN activates computation only in informative regions instead of processing the full feature map densely. Experiments on a three-phase inverter platform show that the proposed method achieves 11 microjoules per diagnosis, corresponding to a 382 times inference-energy reduction compared with a GPU-based CNN, while maintaining 100% diagnostic accuracy. Robustness is further validated under unbalanced loading, current amplitude step changes, and injected measurement noise.
2026-07-27 | Stefan Scholze, Johannes Partzsch, Sebastian Höppner et al.
In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphic hardware has long been advocated as an upcoming alternative to deep networks, taking inspiration from the brain for achieving unprecedented energy efficiency. However, demonstrations of these gains only recently began to grow in complexity and real-world applicability. With SpiNNaker2, we present a chip that bridges the gap between deep networks and neuromorphic computing and allows for flexible exploration of computing approaches that combine both worlds. It features 152 processing elements equipped with an ARM M4F processor and dedicated accelerators, an extended SpiNNaker routing fabric for scalable event-based communication and a range of external interfaces for system integration, including Gbit Ethernet and an LPDDR4 memory interface. We demonstrate performance and efficiency of the SpiNNaker2 chip for neuromorphic and deep network workloads, as well as novel event-based computing approaches. For deep network workloads, the chip achieves up to 4.5 TOPS in high performance mode and up to 2.7 TOPS/W efficiency in high efficiency mode for INT8 workloads. The chip supports spiking neural networks with >150000 neurons and >1.8 billion synaptic events/s when simulated with a 1 ms time step. Its low baseline power of less than 250 mW allows for efficiency even under varying workload conditions, allowing to explore sparse and event-based modes of computation. All this demonstrates the chip's capabilities as a universal hardware platform for scalable brain-inspired computing and its combinations with mainstream deep network approaches.
2026-07-27 | Ansgar Jüngel, Zhiwei Sun, Sara Xhahysa
A structure-preserving fully implicit Scharfetter-Gummel finite-volume scheme for a three-species drift-diffusion model for semiconductors is proposed and analyzed. The equations describe the evolution of the electron, hole, and oxygen vacancy densities in a (bounded) memristor device, coupled to the Poisson equation for the electric potential, with mixed-type boundary conditions. Recasting the Scharfetter-Gummel fluxes in an upwind form, a hidden Fisher information component is revealed. Owing to the degeneracy of the Bernoulli function appearing in the fluxes, additional edgewise coercivity estimates are required, leading to refined local and global dissipation estimates. Using these ideas, the existence of a discrete finite-volume solution, a discrete free energy inequality, and the convergence of the numerical scheme are established. Numerical simulations in two space dimensions confirm the structure-preserving properties of the scheme and illustrate the filament formation in a memristor device.
Autonomous Physical Computation: A Categorical Closure Criterion for Physical and Neuromorphic Reservoirs
2026-07-27 | Nima Dehghani
Physical reservoirs, neuromorphic devices, and wave-mediated systems often possess memory, feedback, and rich state-dependent dynamics, but these properties do not by themselves establish autonomous computation. Here we develop a closure criterion for autonomous physical computation, motivated by the wave--particle walker. We formulate the walker as a stroboscopic reservoir with state, where the wave field stores an exponentially decaying trace of previous droplet impacts and guides future motion through local slope coupling. This model separates physical writing, storage, reading, feedback, and externally triggered erasure. We then define computation as robust coarse-grained transition preservation: a physical map implements an abstract transition only when a coarse-graining satisfies compositionality, with abstract states realized by separated physical basins and transitions stable under noise. Autonomous physical computation requires a further closure condition: an internal physical readout state must select the next physical operation. This criterion classifies the wave--particle walker as a wave-memory machine with genuine Turing-like primitives, but not as a closed autonomous physical computer, because the erasing phase shift is externally imposed. The framework turns this distinction into a design principle: memory becomes autonomous computation when physical readout basins are coupled back to operation selection.
2026-07-26 | Till Zellweger, Marko Mladenović, Kevin Portner et al.
Phase-change memory (PCM) is a mature technology for fast, scalable, non-volatile data storage, with applications spanning embedded memory, as well as in-memory and neuromorphic computing. PCM predominantly relies on chalcogenide alloys, with
$\mathrm{Ge_2Sb_2Te_5}$ (GST) as the industry standard. Yet in these alloys, the individual Ge, Sb, and Te atoms redistribute upon cycling, causing stochastic operation and ultimately device failure. To address this issue, elemental antimony was proposed as a PCM material, but it exhibits a metastable amorphous state that prevents reliable data retention. Moreover, tellurium and antimony can contaminate complementary metal-oxide-semiconductor (CMOS) production lines or act as unintended dopants, restricting manufacturing of PCM to dedicated fabs. Here we introduce elemental germanium (Ge) as a CMOS-native phase-change material that overcomes these fundamental limitations. In a vertical PCM cell architecture, Ge enables sub-nanosecond crystallization (240 ps, 40 times faster than GST), non-volatile data storage with excellent thermal stability ($>$110 °C for 10 years vs. $\sim$87 °C for GST), and a resistance drift coefficient approximately 60% lower than in GST. These results establish pure Ge, a standard semiconductor, as an alternative to chalcogenide phase-change materials, achieving superior performance in key metrics and enabling phase-change memory to be fabricated in standard semiconductor facilities.
2026-07-24 | Tergel Molom-Ochir, Benjamin F. Morris, Yintao He et al.
Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge deployment. In-memory computing (IMC) is energy-efficient on regular workloads but has been considered incompatible with irregular multi-phase algorithms. We introduce phase-to-primitive decomposition, which reformulates each algorithmic phase as a hardware-native IMC primitive. Applied to MCTS, selection, expansion, rollout and backpropagation map to content-addressable memory, combinational logic, a resistive random-access memory (RRAM) crossbar and static random-access memory, keeping search on chip. At 22 nm with fabricated RRAM-array parameters, IMC-MCTS consumes ~60 mW for 9x9 Go, achieving 96x energy efficiency over a central processing unit (CPU) and 65x-2,059x over an H100 graphics processing unit (GPU). It reaches a European Go Federation rating within sample-size uncertainty of open-source Go engines (Pachi-UCT and Michi-C). The same substrate runs eight applications across four AI domains.
2026-07-24 | Mattias Westerink, Sameed Sohail, Berend-Jan van der Zwaag et al.
Event-driven neuromorphic inference exploits activation sparsity by updating neuron state only on spikes. However, weight sparsity introduces irregular gather-style updates that undermine lockstep Single Instruction Multiple Data (SIMD) execution. We call the resulting overheads in control, metadata, and memory activity the sparsity tax. This paper quantifies that tax by comparing three closely related accelerators integrated into one neuromorphic core: (i) baseline lockstep SIMD, (ii) bitmap-gated Sparse-SIMD that selectively disables lanes without compressing weights, and (iii) a Single Instruction Multiple Threads (SIMT) style design with per-PE address generation and run-length coded sparse weights. Using an RTL-to-gates flow in GF22FDX+ and activity-driven energy estimation, we evaluate event-driven neural network inference across varying post-training pruning levels. Results show that total core area changes are insignificant because SRAM dominates area, while performance and energy strongly depend on how sparsity is handled: SIMD and Sparse-SIMD exhibit near-constant throughput, Sparse-SIMD achieves limited energy savings due to bitmap and dense-storage overheads, and SIMT provides the strongest energy scaling and substantial speedups at high sparsity, albeit with sublinear gains due to metadata reads, load imbalance, and sparsity-independent phases. The hardware code for the proposed architectures and experiments is publicly accessible for research purposes.
Neuromorphic Diffusion Language Models: Addressing Compute and Memory Bottlenecks via Sparsity and Block Denoising
2026-07-24 | Dengyu Wu, Clement Ruah, Jiechen Chen et al.
Autoregressive (AR) large language models (LLMs) are inherently inefficient at inference time because each generated token requires accessing the full set of model parameters, leading to low operational intensity and high energy consumption. Masked diffusion language models (MDLMs) partially address this limitation for memory-bound settings by allowing multiple tokens to be generated per parameter access. In order to further enhance inference efficiency on modern platforms with extensive in-chip memory, this work proposes neuromorphic MDLMs (N-MDLMs), which integrate block diffusion with spike-based neuromorphic computation to jointly improve throughput and energy efficiency. While block diffusion increases token throughput by producing multiple tokens per parameter access, spike-induced sparsity reduces effective parameter traffic and computations by skipping inactive channels. To analyze the synergistic effect of sparsity and diffusion, we develop a token-level roofline-inspired model that captures the combined impact of block-parallel generation and spike sparsity on decoding efficiency. Experimental results on translation tasks show that, thanks to spike-induced sparsity, N-MDLMs achieve substantial improvements in energy efficiency and throughput even in compute-bound platforms for which MDLMs would fail to improve over AR-LLMs.
HEMERA: A Heterogeneous Memory-Centric Accelerator with Recursive Dataflow for Edge-Constrained State-Space-Duality Models Inference
2026-07-24 | Hao Ding, Ling Liang, Ruitong Qiao et al.
Structured State Space Models (SSMs), such as Mamba, enable efficient long-sequence modeling with linear time complexity. Recent implementations realize this capability through Structured State Space Duality (SSD), which transforms recursive state evolution into matrix-form computations. However, SSD introduces substantial system-level overheads, including quadratic intermediate materialization, irregular data movement, and prefix-dependent execution, leading to excessive memory traffic and bandwidth demand on conventional architectures. Although prior accelerators mitigate these overheads through optimized dataflows or compute-in-memory techniques, they largely retain matrix-oriented SSD execution and cannot simultaneously avoid quadratic intermediate storage and efficiently map dependency-bound state propagation. This paper presents HEMERA, a heterogeneous memory-centric accelerator for efficient Mamba-2 inference. Rather than directly executing the matrix-form SSD computation, HEMERA reformulates it into an algebraically equivalent streaming-recursive dataflow that avoids quadratic intermediate storage while preserving the original computation. The resulting heterogeneous execution paradigm maps dense linear operations onto in-memory computing units and recursive state updates onto a dedicated streaming engine. Across Mamba-2 models ranging from 130M to 2.8B, HEMERA achieves average latency speedups of 1.4x-3.6x and energy-efficiency improvements of 12.2x-27.0x over the official optimized fused Mamba-2 kernel on NVIDIA A100. It further reduces the average SSD-related execution-time ratio across model scales to 14.12% during long-sequence inference, demonstrating its potential for efficient deployment under edge constraints.
2026-07-24 | Doruk Efe Gökmen, Michel Fruchart, Dmitrii Zendrikov et al.
Computation is the controlled evolution of a state. Asynchronous evolutions, where all parts of the state change in their own time without stopping each other, put this control in jeopardy. It is in fact a mystery how natural processes perform asynchronous computations using many units with no global orchestration. Here we demonstrate how collective computational abilities can emerge in asynchronous many-body systems. The key insight is to split the physical "hardware" underlying the computation into two asymmetrically coupled parts, analogous to position and momentum in a harmonic oscillator. The resulting inertia nudges the evolution of the state so that the asynchronous computation proceeds in the right order. By treating our inertial asynchronous computer as a nonequilibrium material, we map out its phase diagram numerically and analytically using a framework we dub loop dynamical mean-field theory. We experimentally demonstrate our approach using analog spiking neuromorphic chips designed to mimic actual neurons in the brain. In addition, we construct software that can run on asynchronous hardware: we denoise movies whose clean versions were never seen during training, an instantiation of the generalization transition underlying modern machine learning. Our results point to a general strategy for reliable, decentralized computation in energy-constrained settings from dynamics self-assembly to cell differentiation.
2026-07-23 | Matthew Zimmermann, Julia M. Boyle, Hardit Singh et al.
Programmable optical memristors embedded in photonic integrated circuits (PICs) are emerging as an important technology for high-speed optical storage and in-memory optical computing applications. These devices provide multi-level, non-volatile storage of optical phases that can be interrogated at the speed of light, enabling parallel data readout or energy-efficient multiply-accumulate operations in artificial neural networks. However, there remains several outstanding challenges with existing optical memristor technology including durability, material-induced optical losses, large-scale reconfigurability, or fabrication yield for realistic applications. Here we introduce an analog-programmable photonic memristor based on photonic integrated micro-electromechanical (MEMS) cantilevers produced in a CMOS foundry. The memristor consists of low-loss silicon nitride waveguides, requires no additional back-end materials integration, and is all electrically programmed with electrostatic-piezoelectric forces. We demonstrate up to 5-bit phase storage levels, 50 kbit/s programming speeds, strain-assisted non-volatility lifetimes >1 hour, >1 billion cycle endurance, and stress-tested millions of write-read cycles with pseudorandom bit sequences. We further extend the memory lifetime to several days with simple electronic refresh circuits in a portable battery-powered module, demonstrating a proof-of-concept optical random-access memory in static or dynamic configurations. Our MEMS-photonics technology represents an important step toward practical optical memristors.
2026-07-23 | R. Tyumenev, D. S. Kalashnikov, B. V. Fradkin et al.
Superconductor-ferromagnet hybrid structures with tunable kinetic inductance are promising elements for neuromorphic and quantum computing circuits. We report the fabrication and microwave characterization of split-ring resonators based on Nb/Co/Nb/Co/Nb/Al spin-trigger multilayers and demonstrate a non-volatile spin-valve effect on their resonant properties. Reversal of the relative magnetization orientation of the cobalt layers produces a reproducible shift of the resonant frequency up to 4 MHz at zero applied magnetic field, corresponding to a change in the kinetic inductance of the structure. The incorporation of a proximitized aluminum overlayer is shown to enhance the inductance contrast between the parallel and antiparallel magnetic states by a factor of approximately three relative to structures without this layer. The experimental results are in quantitative agreement with a microscopic model based on the Usadel equations. The demonstrated magnetic memory of the resonant frequency at zero field establishes spin-trigger multilayers as viable field-programmable inductive elements for superconducting digital and neuromorphic circuits.
2026-07-30 | Roberto Riaño, Gorka Abad, Stjepan Picek et al.
Backdoor attacks on Spiking Neural Networks (SNNs) have primarily assumed dirty-label poisoning, in which triggered training samples are relabeled to an attacker-selected class. We study clean-label temporal poisoning, where a fixed timestamp transformation is applied only to the target-class training streams, leaving their labels unchanged. The transformation preserves the per-pixel, per-polarity event count exactly, making clean and triggered samples identical after temporal aggregation while altering the sequence processed by the SNN. Across three neuromorphic datasets and both convolutional and transformer-based victims, the attack reaches an ASR of 1.00 in the strongest configurations. We analyze the attack through poison-budget and trigger-shape ablations and evaluate established backdoor defenses adapted to spiking models. Defenses that collapse the time axis before inspection are blind by construction, while feature-space methods detect the poison only in selected settings. Our model-free detector, based on per-step event mass, detects the evaluated temporal transformations, demonstrating both the limitation of rate-collapsed defenses and the boundary of the attack's stealth. To our knowledge, this is the first clean-label backdoor attack evaluated on SNNs and neuromorphic event data.
2026-07-30 | Spyridon Raptis, Haralampos-G. Stratigopoulos
Spiking Neural Networks (SNNs) communicate through sparse binary spike events rather than dense activations, enabling energy-efficient inference on neuromorphic hardware and motivating their use in always-on, battery-powered edge systems. We show that this same efficiency advantage creates a distinct security risk: sponge attacks can increase inference-time spike activity and synaptic workload, inflating energy consumption while remaining difficult to detect through correctness-based monitoring alone. Prior input-space efficiency attacks on SNNs have focused on per-sample optimization, primarily in rate-coded settings. We extend this threat to native event-based binary inputs and study two attack models. First, we develop a per-sample sponge attack that crafts a custom adversarial spike train for each input via gradient-based optimization. This attack increases per-inference SynOps by 1.5-2.6x on three SNN models for the NMNIST, SHD, and IBM DVS Gesture datasets, while preserving the predicted class on at least 98% of evaluated samples. Second, to the best of our knowledge, we introduce the first universal sponge attack for native event-based SNN inputs: a fixed binary perturbation computed offline and applied via XOR to all subsequent inputs. Although weaker, it still inflates SynOps by 1.09-1.24x across all three datasets and represents a more realistic deployment threat because it requires no per-input optimization. Mapping SynOp inflation to estimated Loihi-1 energy yields per-inference overheads from 14 $μ$J to 13.24 mJ. These results show that native event-based SNNs are vulnerable to practical input-space efficiency attacks, and that reusable universal perturbations can accumulate into meaningful battery drain in continuously deployed edge systems.
2026-07-29 | Katharina Bendig, René Schuster, Didier Stricker
Event cameras follow a retina-inspired sensing principle, reporting local intensity changes asynchronously with hightemporal resolution and a wide dynamic range. Spiking Neural Networks (SNNs) complement these sparse event streams through brain-inspired dynamics, using sparse spikes and leaky membrane potentials to integrate information over time. However, many SNN object detectors process isolated event intervals with a single label and reset the network state after each prediction, thereby underusing temporal information in continuous event streams. We introduce Sequence-SOD, a sequence-aware SNN object detector that processes extended event sequences containing labels at multiple time points. Events are accumulated into short intervals, discretized into temporal steps, and fed sequentially to an SSD-style Spiking DenseNet while preserving membrane potentials across intervals within a sequence, so that detection is driven by an evolving neural state instead of independently reset input windows. On the Gen1 Automotive Detection Dataset, sequence-aware training improves mAP from 23.38 for single-interval training to 25.30 without augmentation and to 26.88 withevent-data augmentation. The model achieves a theoretical prediction frequency of 40 Hz. Training and evaluating SNN object detectors on extended event sequences improves their ability to exploit temporal cues while preserving the energy-efficiency benefits of sparse spiking computation. The results highlight sequence-aware training as a complementary direction to architectural improvements for event-based SNN detection.
2026-07-29 | Zeyu Wang
Spiking neural networks (SNNs) are promoted as an energy-efficient substrate because sparse, event-driven activity replaces dense multiply-accumulates with cheap accumulates. We argue the energy dividend of sparsity is not a property of SNNs but of the task. Holding architecture fixed and swapping only the hidden unit (continuous vs. leaky-integrate-and-fire), plus a two-sided target-firing-rate probe, we measure how far activity can be pushed down before quality breaks. Low-load feed-forward perception sparsifies to 5% firing at no accuracy cost; a recurrent language model cannot go below ~50% -- the recurrent state must stay active to carry information. A spiking Transformer, by contrast, sparsifies freely to 2% (3 seeds) -- so the ceiling is a property of recurrent compression, not sequence modeling. Attention escapes the floor only by storing the full key-value cache, trading a firing floor for a memory wall: on neuromorphic hardware, recurrence and attention pay on different axes, neither escapes. We formalize the ceiling with an information-theoretic bound rho >= H_b^{-1}(log2 M / H) and confirm its predictions: the floor rises with memory load, falls with state width, and (refuting a naive memory-only reading) rises with task difficulty. A layer-wise input floor further caps op reduction under dense input, isolating event-driven perception as where neuromorphic hardware wins.
2026-07-29 | Shuhei Ikemoto
A Noise-modulated Neural Network (NNN) learns and infers only in the presence of noise, treating noise as a computational resource rather than a disturbance. The noise lets it learn efficiently by backpropagation while transmitting spike-like signals, but backpropagation needs a reverse path through transposed weights, the weight transport problem, which undermines biological and neuromorphic plausibility. Forward-only alternatives typically substitute a different objective or fixed random feedback, sacrificing stability and accuracy. We show that backpropagation itself can be reconstructed in the NNN from forward-pass statistics alone: a weight mirror estimates each weight matrix from the covariance between a previous-layer unit's output and the next-layer unit's input, and combining it with local differential estimation inside the units propagates the output error recursively along the computational graph, with no transposed-weight readout and no backward data path. The resulting gradient is empirically near-unbiased, and with local per-weight Adam updates it matches the final accuracy of backpropagation on simple regression tasks. With uniformly distributed noise, the local operations reduce to polynomials and comparators, making the whole system, learning rule included, well suited to digital circuits. Thus, in the NNN, noise is a resource not only for inference but also for reconstructing backpropagation.
2026-07-24 | Melissa Lober, Alp Inangu, Gorka Peraza Coppola et al.
Computing centers today mostly operate conventional CPU- and GPU-based systems, where the direct way of decreasing energy consumption is a reduction in the applications' runtime. Neuromorphic computing promises an alternative architecture with improved energy efficiency for artificial intelligence. In this endeavor, code for the simulation of large-scale spiking networks on conventional supercomputers is the reference. We show that turning off automatic NUMA balancing may reduce energy consumption by 30%. This dwarfs other attempts of increasing the energy efficiency of a computing center with respect to cost effectiveness. The memory access pattern of spiking network simulation code dynamically interacts with automatic NUMA balancing. This does not affect the correctness of simulation results and thus goes unnoticed in day-to-day neuroscience research. In performance analysis, however, time measurements fluctuate obstructing attempts to optimize simulation technology. A new time- and compute-node resolved performance display exposes the fine-grained temporal variability of distributed spiking network simulations. The analysis uncovers that automatic NUMA balancing is of disadvantage and affects the jemalloc library for thread-aware memory allocation in a transient manner. The method also allows developers to detect perturbations of the HPC system and target specific improvements to simulation technology. As a consequence, we have equipped our supercomputers with an option to turn on or off automatic NUMA balancing on a per-job basis on the user level. This gives researchers the opportunity to find the best setting for the application at hand. There are indications in the literature that the effect has been observed before, yet it does not seem common knowledge in scientific computing. It remains to be investigated how widespread the phenomenon is among scientific codes.
Neuromorphic Non-Orthogonal Multiple Access for Parallel Remote Inference via Vector Symbolic Architecture
2026-07-24 | Jiechen Chen, Zihang Song, Dengyu Wu et al.
Emerging edge intelligence systems increasingly rely on dense deployments of always-on sensors that must convey task-relevant information to a remote model under tight energy and spectral budgets. The deployment of event-driven neuromorphic sensing paired with spiking neural networks (SNNs) is attractive in this regime because it produces dynamically sparse representations, so that energy is spent on communication and computation only when informative events occur. Prior multiple-access protocols for remote inference using neuromorphic sensing and computing targeted collaborative settings, in which the server fuses information from all devices into a single decision. This paper instead addresses parallel remote inference, in which each device observes a distinct input, and requires its own classification decision. We propose NOMA-NC, a non-orthogonal multiple-access (NOMA) neuromorphic communication (NC) protocol built on the vector symbolic architecture (VSA) framework. In NOMA-NC, each device binds its sparse spike feature map with a device-specific permutation key, and all devices in a group transmit concurrently so that the over-the-air superposition directly realizes the VSA bundling operation. A shared decoding SNN, together with lightweight per-device learned unbinding, recovers all decisions in a single inference pass. Experiments on the N-MNIST and DVS128 Gesture datasets show that NOMA-NC yields goodput gains and savings in terms of receiver computing energy that are sub-proportional to the number of simultaneously active devices, without increasing the per-device transmission energy.
2026-07-22 | Shijie Pan, Agustin Castellano, Zeyu Shen et al.
Learning-enabled decision systems often use offline data or computation to reduce online compute cost. Despite the empirical success of such approaches, there is limited general understanding of how much offline information is needed to achieve a desired accuracy under a fixed online computation budget. We study this question through the lens of amortized parametric optimization: an offline phase stores a finite memory of solved problem instances, and an online phase produces a solution to a new instance by retrieving a warm start and applying
$K$ steps of projected gradient descent. We analyze this setup for smooth convex parametric optimization over a compact domain, using a nonparametric predictor built from the stored offline solutions. For$μ$ -strongly convex objectives, we establish matching upper and lower bounds on the memory required to guarantee$\varepsilon$ -accuracy under a fixed online iteration budget$K$ . For convex objectives satisfying a$β$ -growth condition ($β>2$ ), we obtain near-matching bounds and identify a phase transition in$K$ beyond which additional memory provides no benefit. We further provide a general proof framework that (i) explicitly quantifies the memory cost of acceleration---how much offline memory is required to achieve a prescribed speedup over the unaided online optimizer---and (ii) identifies two key quantities driving this cost: the convergence rate of the online optimizer and the Lipschitz sensitivity of the solution map to the problem parameter. Experiments on parameterized ridge regression confirm the predicted memory--computation--accuracy tradeoffs.
2026-07-29 | B. Neichel, T. Pichon, G. Bourdarot et al.
Future extensions of the VLTI aim to push infrared interferometry toward kilometer scale baselines, enabling angular resolutions of a few tens of micro-arcseconds. A first step would couple a new telescope on the VISTA platform to the existing VLTI infrastructure, creating a 1.4 km baseline at Paranal. Among the possible beam-transport solutions, direct free-space transmission with adaptive-optics (AO) pre-compensation, inspired by Free-Space Optical (FSO) communications, offers an attractive combination of spectral flexibility, moderate cost, and preservation of field information. We present a preliminary AO dimensioning study, including fitting error, photon noise, anisoplanatism, and scintillation, showing that moderate-order correction (around 10x10 actuators) may be sufficient, but that performance depends critically on the poorly known turbulence distribution along the horizontal path. We therefore propose a dedicated turbulence-monitoring experiment across the VISTA-VLTI line of sight, using two 30-50 cm telescopes equipped with calibrated light sources. The experiment combines a wide-field Shack-Hartmann sensor for tomographic reconstruction of the turbulence volume with an event-based (neuromorphic) camera capable of capturing fast, anisotropic turbulence at microsecond timescales. The resulting dataset will provide the first systematic characterization of kilometer-scale horizontal turbulence at Paranal, paving the way toward an extended kilometer-baseline VLTI.
2026-07-28 | Yuhan Bao, Chenxin Shao, Kaiwei Wang
Conventional frame-based Shack--Hartmann wavefront sensors (SHWFS) are limited by dynamic range and the intrinsic trade-off between spatial and temporal resolution, while high-bandwidth acquisition poses additional challenges for real-time wavefront reconstruction. This work presents a real-time, megapixel, kilohertz neuromorphic SHWFS to overcome these limitations. In static optical metrology, the proposed pipeline achieves one-shot wavefront acquisition under extreme illumination non-uniformity, reaching a dynamic range of 260 dB at a 20 Hz acquisition frequency. Owing to this high dynamic range and the concomitant high-intensity resolution, wavefront reconstruction errors in dim and bright sub-apertures are reduced by 59% and 70%, respectively, relative to conventional frame-based SHWFS. For dynamic wavefront sensing, the system provides kilohertz-rate centroid tracking over a megapixel field of view with microsecond-scale latency. Centroid localization errors are 0.18 pixels during optical alignment supervision and 0.26 pixels in high-speed turbulence observation, verifying the accuracy and reliability of the system across dynamic scenarios. The per-sub-aperture processing throughput reaches 420,737 Hz on a standard CPU, demonstrating high-speed real-time computation without specialized hardware acceleration. Together, these results establish a unified neuromorphic SHWFS framework for high-fidelity one-shot static wavefront reconstruction and real-time high-bandwidth dynamic wavefront sensing.
2026-07-26 | Jianing Li, Dianze Li, Arren Glover et al.
Conventional frame-based cameras face significant challenges in detecting objects under high-speed motion blur or in low-light environments. Neuromorphic cameras provide asynchronous visual streams with high temporal resolution and a wide dynamic range, offering a promising solution for object detection under challenging conditions. Despite the development of numerous models and the emergence of various applications in neuromorphic object detection, there is still a lack of deep understanding and standardized benchmarks to assess progress and address key challenges. In this paper, we provide a comprehensive survey and benchmark of existing neuromorphic object detection algorithms. Specifically, we first present a problem description, review the available datasets, and revisit the evaluation metrics. We then explore existing neuromorphic object detection approaches from various perspectives, including event representation, temporal modeling, multimodal fusion, asynchronous processing, low-latency processing, and energy-efficient computing. Furthermore, we evaluate a wide range of representative neuromorphic object detection models and offer detailed analyses of the comparative results. Finally, we discuss unresolved issues in neuromorphic object detection and propose potential future research directions. We hope this survey and benchmark will be a valuable resource for researchers and provide guidance for future advancements in neuromorphic object detection.
Practical Post-Quantum Cryptography for Bandwidth Constrained or Non-Terrestrial Networks, and Power Constrained Devices
2026-07-25 | Elliot Eichen, Sylvia Llosa, Yueqi Chen et al.
Post-quantum (PQ) cryptographic algorithms, particularly for authentication, are more complex than classical algorithms and require larger certificates, signatures, and keys. Establishing a PQ-secure network connection increases bandwidth, memory, computation time, and energy consumption. These costs are especially severe in Non-Terrestrial Networks (NTNs), where long propagation delays, intermittent connectivity, constrained link budgets, satellite handovers, limited terminal resources, and bandwidth-constrained satellite-to-ground links amplify the overhead of certificate-based PQ authentication. Consequently, applications such as key rotation and key management may be unable to achieve acceptable handshake reliability or support NIST PQ Security Categories above Category 1. Similar limitations affect low-power IoT devices and bandwidth- or energy-constrained terrestrial networks, where PQ authentication may restrict devices to Category 1 security or prevent ambient-powered endpoints from supporting PQ authentication altogether. This paper investigates an alternative cryptographic framework that replaces PQ digital certificates with shared secret keys (SSKs). The framework leverages shared-secret ecosystems that do not rely on asymmetric key distribution, such as 5G/6G, and combines a Key Distribution Center (KDC) (e.g., Kerberos) with a preshared-key PQ handshake (e.g., DTLS-PSK) and ephemeral PQ key establishment (e.g., ML-KEM). Compared with certificate-based PQ authentication (e.g., ML-DSA), the proposed approach reduces handshake bandwidth, endpoint RAM, computation time, and energy use while preserving PQ-secure AEAD, including forward secrecy and replay resistance. Applications include NTN-based key rotation and management, uncrewed aerial vehicle (UAV) command-and-control systems, embedded medical sensors, and supply-chain monitoring and asset-tracking platforms.
2026-07-30 | Jonas Mensing, Wilfred G. van der Wiel, Andreas Heuer
Physical computing leverages complex dynamical systems for energy-efficient data processing. In this work, we present a neuromorphic architecture based on metallic nanoparticles interconnected by molecular junctions on a
$\text{SiO}_2$ /Si substrate. We demonstrate that surrounding static control electrodes transform this nanoparticle network from a passive reservoir into a tunable nonlinear dynamical system. By analyzing how these electrodes route simple one-dimensional voltage inputs into multidimensional signal responses, we establish three core design rules to maximize computational performance. First, operating near the system's cutoff frequency achieves an optimal balance between nonlinear charge tunneling and linear capacitive memory. Second, tuning the underlying$\text{SiO}_2$ thickness sets the electrostatic screening length and dictates the memory type. Thick oxide layers reduce the screening length, causing networks larger than this length to transition into a persistent, non-volatile-like regime. Conversely, networks smaller than the screening length exhibit only fading memory. Third, introducing structural disorder via heterogeneous molecular junctions overcomes inherent limits on expressivity. While a network's computational expressivity scales with its physical size, it is ultimately capped by the screening length. Breaking internal spatial symmetries with localized disorder bypasses this saturation, allowing control voltages to independently manipulate specific signal amplitudes and phases, universally maximizing performance for dynamic neuromorphic applications.