What we work on
Post-Moore Exascale HPC
Our research contributes to the advancement and interaction of HPC and AI. We investigate hardware, suitable algorithms, the application layer, and the surrounding ecosystem. Part of our job is the screening of the HPC/AI landscape for new, emerging solutions. To mitigate risk and discover promising innovations, we evaluate the hardware, software, and the companies developing them. Overall, our tech scouting and research help HLRS and the community deploy the fastest computing systems.
Algorithms, Peak Performance & Optimization
Hardware can only be evaluated meaningfully in combination with a problem statement and a suitable algorithm to solve it. We are interested in the maximum achievable performance of such hardware-algorithm combinations. We actively develop numerical kernels and port them to new hardware, enabling comparisons of runtime costs — such as time-to-solution, performance/$, and performance/watt — across different hardware concepts. Through publications, conference engagement, and community outreach, we inform the community about highly efficient implementations as well as the dead ends our research uncovers.
Programming Models, Libraries & Usability
Peak performance alone rarely drives hardware adoption, usability plays an equally important role for users of any hardware stack. We investigate the usability of emerging hardware and report on both the good examples and the shortcomings of individual concepts. We actively (co)develop libraries that ease hardware usage for our lab and for other investigators. Through open-source software, panel discussions, and research collaborations we give back to the wider HPC and AI community.
Example Hardware Concepts We Investigate
-
SpiNNaker2 (SpiNNcloud) — A brain-inspired neuromorphic chip promising highly performant computation with extremely low energy consumption on sparse networks and other sparse workloads.
-
Wafer-Scale Engine 3 (Cerebras) — A massively parallel AI chip built from an entire wafer, featuring a huge number of cores with small but very fast local memory and an on-chip interconnect.
-
Photonic Chips — We cooperate with several analog photonic chip makers and are currently preparing to investigate their respective chips.
-
RISC-V Chips — We cooperate with several RISC-V HPC chip makers and are currently preparing to investigate their respective chips.
-
Coarse-Grained Reconfigurable Arrays (CGRAs) — Architectures that reconfigure at the level of functional units rather than individual gates, promising near-ASIC efficiency with software-like flexibility for data-flow-oriented workloads.
-
Other Accelerator Concepts — We keep an eye on further emerging designs such as the stencil and tensor accelerator STX by Fraunhofer, and evaluate promising candidates as they mature.