I'm Johann Schmidt.

I'm a Research Scientist.

I'm a Developer.

I'm a Father of Two.

Welcome

based in Magdeburg, Germany.
Open to research lead roles in industry from 2027.

LinkedIn

About Me

Let me introduce myself

I'm Johann Schmidt,
Research Scientist in Deep Learning

Since 2020 I have worked full time as a research scientist at the Artificial Intelligence Lab (AILab) of Otto von Guericke University Magdeburg (OVGU), Germany. I build deep learning systems in collaboration with universities, research institutes and industry partners. These projects range from production scheduling to traffic signal control and acoustic traffic sensing.

I pursue my PhD alongside this work. A strong vision model can change its answer once the view of an object changes. I develop small modules that sit in front of a frozen pretrained model and undo such transformations, which makes the model more robust to them.

I am looking for a research lead position in industry. I will submit my thesis at the end of 2026 and am open to new roles from then on. My current funding ends in April 2027. I am a father of two, so a family-friendly employer matters to me. I speak German (native) and English (daily working language).

Career

Resume

Magdeburg-Stendal University of Applied Sciences, Germany

While I enjoyed designing electrical circuits, I discovered my real passion in computer science. It gave me the power to bring those static circuits to life, and it still motivates me today.

Internships and Student Jobs 2013 - 2016

TCS AG Genthin, Bachmann GmbH Magdeburg, Enercon GmbH Magdeburg

I first came into contact with the development and assembly of electrical circuit boards. I analyzed and evaluated wind turbine data and helped with the cabling of wind turbine control cabinets.

Bachelor Thesis 2017 - 2018

Institute for Automation and Communication (ifak) e.V. Magdeburg, Germany

I developed a GUI in C for a clamp-on ultrasonic sensor system on an embedded controller, and analyzed and evaluated the overall measurement system.

Otto von Guericke University Magdeburg, Germany

I delved into advanced computer science courses, dedicating countless hours to self-study to bridge gaps in my knowledge. I truly found my passion when I dove into artificial intelligence, specifically deep learning.

Master Thesis 2019 - 2020

Fraunhofer Institute for Factory Operation and Automation IFF, Magdeburg, Germany

ETFA 2021

I developed and integrated a vision-based deep learning model for real-time human action recognition, used for human-machine interaction in an industrial environment.

AILab, Otto von Guericke University Magdeburg, Germany

Full-time position on funded R&D projects with industry partners. These projects span a wide range of topics, from production scheduling to traffic signal control and acoustic sensing. This diversity has given me a broad knowledge base. I pursue my PhD alongside this work.

SENECA 2020 - 2022

BMBF-funded project with Thorsis Technologies GmbH and TECTRON GmbH

IFAC INCOM 2021 ESWA 2021

Real-time decision support for production scheduling at the electronics manufacturer Tectron, an NP-hard problem with up to 200 open orders per planning run. I developed deep learning approaches that find solution candidates with larger gains than existing heuristics. Not everything worked: a permutation-invariant network trained on optimal schedules solved small instances but did not scale. The final system therefore lets the planner choose among several algorithms, including reinforcement learning and a genetic algorithm. In Tectron's production, it beat manual planning on late orders and total tardiness.

AI-Engineering 2022

BMBF-funded project of five universities in Saxony-Anhalt

A new Bachelor's program that combines AI with engineering, built jointly by five universities. I contributed to the curriculum in the first project year and led the development of the internal platform for open educational resources, which lets students and teachers share course material across the universities. Students learn through team projects with regional companies and open course material instead of classic lectures. The program has been running since 2023.

Sustainable Supply Chain DeepHack 2022

2nd place, Transatlantic AI Hackathon

In one weekend our team built a system that estimates the volume of parcels automatically and uses it to pack delivery vehicles more efficiently. The route is planned together with the loading order. Parcels are loaded first in, last out, so the next delivery is always within reach.

PASCAL 2022 - 2025

BMBF-funded project with Thorsis Technologies GmbH

AAAI: MALTA 2025

AI-based traffic signal control on real-time V2X data, built for Magdeburg's urban testbed. I developed the core component, a single reinforcement learning policy on a hierarchical graph neural network that controls any intersection layout. Trained only on synthetic road networks, it transfers zero-shot to real ones (Cologne, Ingolstadt). It outperforms heuristic and learned baselines on travel time, waiting time and CO₂ emissions. TransferLight was integrated with Thorsis into an operator-assistance prototype running on edge hardware.

KIVA-NET 2025 - 2027

ZIM-funded project with Thorsis Technologies GmbH

Follow-up to PASCAL with the same industry partner. We are building a low-cost acoustic sensor that counts and classifies vehicles and rates traffic flow from sound alone, in real time on edge hardware. A video tracker labels the audio automatically and has labeled more than 16,000 vehicle passes so far. On this data I benchmark models from compact CNNs to pretrained audio transformers to size the model for the edge device. First models detect passing vehicles with 96.5% accuracy.

Otto von Guericke University Magdeburg, Germany

Thesis research since 2022, alongside my full-time project work. It began with the work on radial beam sampling. I will submit the thesis at the end of 2026.

OxML Summer School August 2025

Oxford, United Kingdom

I attended the Oxford Machine Learning Summer School (OxML) on Representation Learning.

PhD Thesis

AILab, Otto von Guericke University Magdeburg, Germany

Robust and Efficient Discriminative Deep Learning by Canonicalizing and Lifting Symmetry-Degenerated Signals. A strong vision model can change its answer once the view of an object changes. My thesis develops canonicalization modules as a remedy. A small module in front of a frozen pretrained model undoes the transformation of each input, so robustness becomes an add-on that does not require retraining the downstream model. Read the summary.

Research

PhD Thesis

Robust and Efficient Discriminative Deep Learning by Canonicalizing and Lifting Symmetry-Degenerated Signals

Otto von Guericke University Magdeburg · submission planned for the end of 2026

Vision models fail on inputs a person finds trivial. Rotate or rescale an image and a strong classifier can change its answer. In my experiments a ResNet-50 falls from 66% to 39% accuracy on ImageNet-200 once the test images are rotated. The standard fixes are expensive. Retraining with augmented data cost that model five points on upright images. Symmetry-aware architectures give guarantees but have to be designed for each transformation and trained from scratch. Neither helps a team that already owns a large pretrained model.

My thesis develops canonicalization as a third route. A small module in front of the model undoes the transformation of each input and hands over a familiar view. People do the same when they step closer to see an object of interest more sharply. The model stays frozen, so robustness becomes an add-on and requires no retraining or fine-tuning of the downstream model.

Main contributions

A rotation canonicalizer that is capable and cheap

EurIPS: UPLB 2025 Under review

A canonicalizer is usually trained with a prior that treats every training image as upright. Real datasets break this alignment assumption. My bootstrapping algorithm re-aligns the training samples step by step, with convergence guarantees under mild conditions for arbitrary compact groups. On four fine-grained benchmarks it outperforms equivariant and canonicalization baselines and performs on par with augmentation. A second approach replaces the prior itself. I precompute the orientations at which a pretrained model errs least and distill them into a small rotation-equivariant network. The large model never enters the training loop. The same ResNet-50 recovers from 39% to 63% on rotated images and loses one point on upright ones.

Robustness at test time with no training

ICML 2024 ECML-PKDD 2026

Inverse Transformation Search shows a frozen model several transformed versions of an input and keeps the one it finds most familiar. To my knowledge it was the first canonicalizer that needs no training. A follow-up treats the problem as out-of-distribution detection and lifts accuracy on transformed MNIST from 38% to 90%. A gate corrects only inputs that look unfamiliar, which preserves most of the accuracy on normal ones.

Stable spatial transformers and saccadic canonicalization

ECCV: BISCUIT 2026 BMVC 2026

Learned image canonicalization without any inductive biases is brittle and likely to collapse during training. I split the transformation into bounded primitives such as rotation and scale, and prove that the bounds rule out degenerate warps. The module shares weights with a vision transformer and was evaluated on biodiversity and medical imaging data. With translation and scale alone, the same idea yields a multi-view spatial transformer that zooms in like a saccade. When a label-aware oracle picks a single crop, a pretrained ConvNeXtV2-L reaches 98% on ImageNet. The features are there but the model looks in the wrong place. I build a saccader that learns to focus using a small priority head, which selects a few high-resolution regions. One attention block fuses the crops with the global view. This improves on the plain backbone across nine fine-grained benchmarks with few exceptions, and memory no longer grows with input resolution.

Symmetries in production scheduling

IFAC INCOM 2021 HICSS 2024

Symmetries also shape problems outside vision. In job scheduling the open jobs form a set, and parallel machines can be swapped without changing the cost of a schedule. Many optimal schedules are therefore equivalent, which leaves a learning model without a unique target. I encode the jobs with permutation-invariant layers and sort the machines into a canonical order. In a later simulated annealing heuristic, breaking the symmetry between tardy and early jobs lifts degenerate optima and yields schedules in real time.

Symmetries in traffic signal control

AAAI: MALTA 2025

Traffic signal control has the same structure. The lanes, movements and signal phases of an intersection are sets with no natural order. I model them as abstraction levels of a hierarchical graph neural network with permutation-invariant aggregation, so a single policy controls any intersection layout. Where order matters, the symmetry is broken on purpose. A positional encoding relative to the center of the intersection tells the model where along a lane the vehicles are. Trained only on synthetic road networks, the policy transfers zero-shot to real ones.




Papers

Publications

First page of Human Action Recognition as part of a Natural Machine Operation Framework

2021 ETFA

Human Action Recognition as part of a Natural Machine Operation Framework

S. Bexten, J. Schmidt, C. Walter, and N. Elkmann

IEEE International Conference on Emerging Technologies and Factory Automation

Read paper Code

The reliability of systems that use machine learning to recognize the human working in an industrial environment is of high importance for employee safety. We present a framework which is capable of recognizing the person's natural interaction with an industrial machine. We focus on the application of human action recognition in the context of machine operation by skilled workers in industrial or commercial environments. We propose a framework that includes action recognition as part of a software component for understanding behavior. For our use case, we defined an exemplary machine operation workflow which we use to compare five different neural networks in terms of prediction accuracy and real-time capabilities. Moreover, we compare different input shapes, such as the resolution of input images and the size of the possible 3D-volume in order to study the robustness of the models. For our evaluation, we created our own custom dataset containing six action classes. Our analysis shows that the best model is the I3D with color images, a resolution of 112 x 112 pixels and 16 consecutive frames. The I3D also exhibited the best run-time performance for real-time applications.

@inproceedings{Bexten2021,
  title     = {Human Action Recognition as part of a Natural Machine Operation Framework},
  author    = {Bexten, Simone and Schmidt, Johann and Walter, Christoph and Elkmann, Norbert},
  booktitle = {IEEE International Conference on Emerging Technologies and Factory Automation},
  year      = {2021},
  doi       = {10.1109/ETFA45728.2021.9613331},
  url       = {https://ieeexplore.ieee.org/document/9613331}
}
First page of NeuroEvolution of augmenting topologies for solving a two-stage hybrid flow shop scheduling problem: A comparison of different solution strategies

2021 ESWA

NeuroEvolution of augmenting topologies for solving a two-stage hybrid flow shop scheduling problem: A comparison of different solution strategies

S. Lang, T. Reggelin, J. Schmidt, M. Müller, and A. Nahhas

Expert Systems with Applications

Read paper Code

The article investigates the application of NeuroEvolution of Augmenting Topologies (NEAT) to generate and parameterize artificial neural networks (ANN) on determining allocation and sequencing decisions in a two-stage hybrid flow shop scheduling environment with family setup times. NEAT is a machine-learning and neural architecture search algorithm, which generates both the structure and the hyper-parameters of an ANN. Our experiments show that NEAT can compete with state-of-the-art approaches in terms of solution quality and outperforms them regarding computational efficiency. The main contributions of this article are: (i) A comparison of five different strategies, evaluated with 14 different experiments, on how ANNs can be applied for solving allocation and sequencing problems in a hybrid flow shop environment, (ii) a comparison of the best identified NEAT strategy with traditional heuristic and metaheuristic approaches concerning solution quality and computational efficiency.

@article{Lang2021,
  title   = {NeuroEvolution of augmenting topologies for solving a two-stage hybrid flow shop scheduling problem: A comparison of different solution strategies},
  author  = {Lang, Sebastian and Reggelin, Tobias and Schmidt, Johann and M{\"u}ller, Marcel and Nahhas, Abdulrahman},
  journal = {Expert Systems with Applications},
  year    = {2021},
  url     = {https://www.sciencedirect.com/science/article/pii/S095741742100107X},
  doi     = {10.1016/j.eswa.2021.114666},
  volume  = {172},
  pages   = {114666}
}
First page of Approaching Scheduling Problems via a Deep Hybrid Greedy Model and Supervised Learning

2021 IFAC INCOM

Approaching Scheduling Problems via a Deep Hybrid Greedy Model and Supervised Learning

J. Schmidt and S. Stober

IFAC Symposium on Information Control Problems in Manufacturing

Read paper Code

Scheduling still constitutes a challenging problem, especially for complex problem settings involving due dates and sequence-dependent setups. The majority of existing approaches use heuristics or meta-heuristics, like Genetic Algorithms or Reinforcement Learning. We show that a supervised learning framework can learn and generalize from generated optimal target schedules, which amplifies convergence compared to unsupervised methods. We present a deep hybrid greedy framework, which can predict near-optimal schedules by utilizing the following key mechanisms: (i) Through the interplay between heuristics and a deep neural network our hybrid model can combine the benefits. Specifically, complex patterns from optimal schedules can be learned by a neural network. We reduce the computational costs by outsourcing trivial decisions to heuristics, thereby allowing consistent decisions during training. (ii) The problem complexity can be reduced by employing a greedy prediction scheme, where one job at a time is predicted. (iii) We propose a re-scheduling mechanism for idle jobs, which enables long-term cost reduction and renders the framework reactive and dynamic. Through the heuristics and the neural network, our model is real-time capable during inference. We compare our model against prevailing scheduling heuristics and our model outperformed one of them in terms of makespan and lateness minimization. The key purpose of this work is to give a proof of concept that supervised learning is applicable for complex scheduling problems.

@inproceedings{Schmidt2021,
  title     = {Approaching Scheduling Problems via a Deep Hybrid Greedy Model and Supervised Learning},
  author    = {Schmidt, Johann and Stober, Sebastian},
  booktitle = {IFAC Symposium on Information Control Problems in Manufacturing},
  year      = {2021},
  url       = {https://www.sciencedirect.com/science/article/pii/S2405896321008417},
  doi       = {https://doi.org/10.1016/j.ifacol.2021.08.095},
  pages     = {805--810},
  volume    = {54}
}
First page of Learning Continuous Rotation Canonicalization with Radial Beam Sampling

2023 ArXiv

Learning Continuous Rotation Canonicalization with Radial Beam Sampling

J. Schmidt and S. Stober

Read paper Code

Nearly all state-of-the-art vision models are sensitive to image rotations. Existing methods often compensate for missing inductive biases by using augmented training data to learn pseudo-invariances. Alongside the resource-demanding data inflation process, predictions often poorly generalize. The inductive biases inherent to convolutional neural networks allow for translation equivariance through kernels acting parallel to the horizontal and vertical axes of the pixel grid. This inductive bias, however, does not allow for rotation equivariance. We propose a radial beam sampling strategy along with radial kernels operating on these beams to inherently incorporate center-rotation covariance. Together with an angle distance loss, we present a radial beam-based image canonicalization model, short BIC. Our model allows for maximal continuous angle regression and canonicalizes arbitrary center-rotated input images. As a pre-processing model, this enables rotation-invariant vision pipelines with model-agnostic rotation-sensitive downstream predictions. We show that our end-to-end trained angle regressor is able to predict continuous rotation angles on several vision datasets, namely FashionMNIST, CIFAR10, COIL100, and LFW.

@misc{Schmidt2023,
  title  = {Learning Continuous Rotation Canonicalization with Radial Beam Sampling},
  author = {Schmidt, Johann and Stober, Sebastian},
  year   = {2023},
  url    = {https://arxiv.org/abs/2206.10690},
  doi    = {10.48550/arXiv.2206.10690}
}
First page of Reviving Simulated Annealing: Lifting its Degeneracies for Real-Time Job Scheduling

2024 HICSS

Reviving Simulated Annealing: Lifting its Degeneracies for Real-Time Job Scheduling

J. Schmidt, B. Köhler, and H. Borstell

Hawaii International Conference on System Sciences (HICSS)

Read paper Code

Inspired by the success of Simulated Annealing in physics, we transfer insights and adaptations to the scheduling domain, specifically addressing the one-stage job scheduling problem with an arbitrary number of parallel machines. In optimization, challenges arise from local optima, plateaus in the loss surface, and computationally complex Hamiltonian (cost) functions. To overcome these issues, we propose the integration of corrective actions, including symmetry breaking, restarts, and freezing out non-optimal fluctuations, into the Metropolis-Hastings algorithm. Additionally, we introduce a generalized Hamiltonian that efficiently fuses straightforward but widely applied processing-time cost functions. Our approach outperforms decision rules, meta-heuristics, and novel reinforcement learning algorithms. Notably, our method achieves these superior results in real-time, thanks to its computationally efficient evaluation of the Hamiltonian.

@inproceedings{Schmidt2024a,
  title     = {Reviving Simulated Annealing: Lifting its Degeneracies for Real-Time Job Scheduling},
  author    = {Schmidt, Johann and K{\"o}hler, B. and Borstell, H.},
  booktitle = {Hawaii International Conference on System Sciences (HICSS)},
  year      = {2024},
  url       = {https://aisel.aisnet.org/hicss-57/da/digital_twins/3/},
  doi       = {10.24251/HICSS.2024.208}
}
First page of Pursuing the Perfect Projection: A Projection Pursuit Framework for Deep Learning

2024 WSOM+

Pursuing the Perfect Projection: A Projection Pursuit Framework for Deep Learning

J. Perschewski, J. Schmidt, and S. Stober

International Workshop on Self-Organizing Maps and Learning Vector Quantization, Clustering and Data Visualization (WSOM+)

Read paper Code

The curse of dimensionality refers to phenomena occurring with increasing dimensionality such as marginal differences in distances. Projection pursuit solves this issue by projecting high-dimensional data into a low-dimensional space where meaningful distances allow unbiased function estimation. However, projection pursuit only considers projections onto lines and the unbiased function depends on the sample size. We introduce deep projection pursuit (DPP) to remedy these limitations by using an ensemble of projections on parameterized surfaces combined with neural networks to solve the learning task. Furthermore, we demonstrate the capabilities of the DPP framework by training principal component curves and solving supervised tasks with interpretable models. Finally, we show the ability to maintain group properties in the projection space. Due to these applications, deep projection pursuit is a flexible design paradigm with various use cases.

@inproceedings{Perschewski2024,
  title     = {Pursuing the Perfect Projection: A Projection Pursuit Framework for Deep Learning},
  author    = {Perschewski, J. and Schmidt, Johann and Stober, Sebastian},
  booktitle = {International Workshop on Self-Organizing Maps and Learning Vector Quantization, Clustering and Data Visualization (WSOM+)},
  year      = {2024},
  url       = {https://link.springer.com/chapter/10.1007/978-3-031-67159-3_6},
  doi       = {10.1007/978-3-031-67159-3_6},
  pages     = {43--52}
}
First page of Tilt your Head: Activating the Hidden Spatial-Invariance of Classifiers

2024 ICML

Tilt your Head: Activating the Hidden Spatial-Invariance of Classifiers

J. Schmidt and S. Stober

International Conference on Machine Learning (ICML)

Read paper Code

Deep neural networks are applied in more and more areas of everyday life. However, they still lack essential abilities, such as robustly dealing with spatially transformed input signals. Approaches to mitigate this severe robustness issue are limited to two pathways: Either models are implicitly regularized by increased sample variability (data augmentation) or explicitly constrained by hard-coded inductive biases. The limiting factor of the former is the size of the data space, which renders sufficient sample coverage intractable. The latter is limited by the engineering effort required to develop such inductive biases for every possible scenario. Instead, we take inspiration from human behavior, where percepts are modified by mental or physical actions during inference. We propose a novel technique to emulate such an inference process. This is achieved by traversing a sparsified inverse transformation tree during inference using parallel energy-based evaluations. Our proposed inference algorithm, called Inverse Transformation Search (ITS), is model-agnostic and equips the model with zero-shot pseudo-invariance to spatially transformed inputs. We evaluated our method on several benchmark datasets, including a synthesized ImageNet test set. ITS outperforms the utilized baselines on all zero-shot test scenarios.

@inproceedings{Schmidt2024b,
  title     = {Tilt your Head: Activating the Hidden Spatial-Invariance of Classifiers},
  author    = {Schmidt, Johann and Stober, Sebastian},
  booktitle = {International Conference on Machine Learning (ICML)},
  volume    = {235},
  pages     = {43705--43722},
  year      = {2024},
  url       = {https://proceedings.mlr.press/v235/schmidt24a.html}
}
First page of TransferLight: Zero-Shot Traffic Signal Control on any Road-Network

2025 AAAI: MALTA

TransferLight: Zero-Shot Traffic Signal Control on any Road-Network

J. Schmidt, F. Dreyer, S. A. Hashimi, and S. Stober

Association for the Advancement of Artificial Intelligence: Multi-Agent Reinforcement Learning for Transportation Autonomy Workshop

Read paper Code

Traffic signal control plays a crucial role in urban mobility. However, existing methods often struggle to generalize beyond their training environments to unseen scenarios with varying traffic dynamics. We present TransferLight, a novel framework designed for robust generalization across road networks, diverse traffic conditions and intersection geometries. At its core, we propose a log-distance reward function, offering spatially aware signal prioritization while remaining adaptable to varied lane configurations, overcoming the limitations of traditional pressure-based rewards. Our hierarchical, heterogeneous, and directed graph neural network architecture effectively captures granular traffic dynamics, enabling transferability to arbitrary intersection layouts. Using a decentralized multi-agent approach, global rewards, and novel state transition priors, we develop a single, weight-tied policy that scales zero-shot to any road network without re-training. Through domain randomization during training, we additionally enhance generalization capabilities. Experimental results validate TransferLight's superior performance in unseen scenarios, advancing practical, generalizable intelligent transportation systems to meet evolving urban traffic demands.

@inproceedings{Schmidt2025a,
  title     = {TransferLight: Zero-Shot Traffic Signal Control on any Road-Network},
  author    = {Johann Schmidt and Frank Dreyer and Sayed Abid Hashimi and Sebastian Stober},
  booktitle = {Association for the Advancement of Artificial Intelligence: Multi-Agent Reinforcement Learning for Transportation Autonomy Workshop},
  year      = {2025},
  doi       = {10.48550/arXiv.2412.09719},
  url       = {https://arxiv.org/abs/2412.09719}
}
First page of Robust Canonicalization through Bootstrapped Data Re-Alignment

2025 EurIPS: UPLB

Robust Canonicalization through Bootstrapped Data Re-Alignment

J. Schmidt and S. Stober

Conference on Neural Information Processing Systems (European Satellite): Unifying Perspectives on Learning Biases Workshop

Read paper Code

Fine-grained visual classification (FGVC) tasks, such as insect and bird identification, demand sensitivity to subtle visual cues while remaining robust to spatial transformations. A key challenge is handling geometric biases and noise, such as different orientations and scales of objects. Existing remedies rely on heavy data augmentation, which demands powerful models, or on equivariant architectures, which constrain expressivity and add cost. Canonicalization offers an alternative by shielding such biases from the downstream model. In practice, such functions are often obtained using canonicalization priors, which assume aligned training data. Unfortunately, real-world datasets never fulfill this assumption, causing the obtained canonicalizer to be brittle. We propose a bootstrapping algorithm that iteratively re-aligns training samples by progressively reducing variance and recovering the alignment assumption. We establish convergence guarantees under mild conditions for arbitrary compact groups, and show on four FGVC benchmarks that our method consistently outperforms equivariant and canonicalization baselines while performing on par with augmentation.

@inproceedings{Schmidt2025b,
  title     = {Robust Canonicalization through Bootstrapped Data Re-Alignment},
  author    = {Johann Schmidt and Sebastian Stober},
  booktitle = {Conference on Neural Information Processing Systems (European Satellite): Unifying Perspectives on Learning Biases Workshop},
  year      = {2025},
  doi       = {10.48550/arXiv.2510.08178},
  url       = {https://arxiv.org/abs/2510.08178}
}
First page of Geometrically Constrained and Token-Based Probabilistic Spatial Transformers

2026 ECCV: BISCUIT

Geometrically Constrained and Token-Based Probabilistic Spatial Transformers

J. Schmidt, T. Siegl, M. Becker, and S. Stober

European Conference on Computer Vision: International Workshop on Biomedical Image and Signal Computing for Unbiasedness, Interpretability, and Trustworthiness

Read paper Code

Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification. Careful consideration is required for reliable use in high-stakes settings. A model should stay robust under such transformations, expose why a correction was applied, and signal when its input is ambiguous. While geometrically equivariant architectures provide a mathematically grounded solution, they often limit model flexibility through strict symmetry constraints and incur significant computational overhead. Spatial Transformer Networks (STNs) offer a data-driven, flexible alternative for learning pseudo-equivariances to affine transformations. However, STNs have historically been restricted to convolutional architectures and suffer from training instability. To address this, we introduce a novel STN framework. It leverages the global modeling capabilities of transformers to regress the affine transformation acting on the input. For this, we decompose affine transformations into interpretable primitives, regressed under adaptable geometric constraints, thereby preventing the training instability typically caused by degenerate transformations. By sharing weights between the localization network and the classification backbone, the framework requires minimal computational overhead. Extensive experiments on challenging insect biodiversity and medical imaging benchmarks demonstrate that our approach achieves superior predictive performance under diverse spatial transformations while maintaining high efficiency.

@inproceedings{Schmidt2026a,
  title     = {Geometrically Constrained and Token-Based Probabilistic Spatial Transformers},
  author    = {Schmidt, Johann and Siegl, Tom and Becker, Martin and Stober, Sebastian},
  booktitle = {European Conference on Computer Vision: International Workshop on Biomedical Image and Signal Computing for Unbiasedness, Interpretability, and Trustworthiness},
  year      = {2026},
  url       = {https://openreview.net/forum?id=8C4vVAbMW8}
}
First page of PPS: Plug-and-Play Saccadic Vision for Fine-Grained Classification

2026 BMVC

PPS: Plug-and-Play Saccadic Vision for Fine-Grained Classification

J. Schmidt, S. Stober, J. Denzler, and P. Bodesheim

British Machine Vision Conference

Read paper Code

Fine-grained visual classification benefits from high-resolution inputs to resolve subtle, localized cues, but at the cost of large compute and memory budgets. Inspired by human saccadic vision, we propose a lightweight extension that turns any pretrained backbone with spatial feature maps into a saccadic classifier through glimpse-induced inference. The only architectural additions are a small priority head and a single cross-attention fusion block. The same weight-tied encoder encodes a downsampled peripheral glance, exposes the spatial feature map from which the priority head predicts a glimpse-priority distribution and re-encodes high-resolution glimpse crops sampled from that distribution in parallel. A sequential sampler with non-maximum suppression draws a runtime-configurable number of glimpses. A residual multi-head cross-attention block then fuses the glance against the glimpses. To improve the quality of the priority maps, we train the priority head by a per-glimpse information-gain target from a lightweight per-image memory bank. This requires no reinforcement learning. Complexity is linear in the number of glimpses and peak memory is decoupled from input resolution. Across nine benchmarks our framework improves consistently over its backbones.

@inproceedings{Schmidt2026b,
  author    = {Johann Schmidt and Sebastian Stober and Joachim Denzler and Paul Bodesheim},
  title     = {PPS: Plug-and-Play Saccadic Vision for Fine-Grained Classification},
  year      = {2026},
  booktitle = {British Machine Vision Conference},
  url       = {https://arxiv.org/abs/2509.15688}
}
First page of Zero-Shot Test-Time Canonicalization using Out-of-Distribution Scoring

2026 ECML-PKDD

Zero-Shot Test-Time Canonicalization using Out-of-Distribution Scoring

D. Lindner, J. Schmidt, T. Siegl, M. Becker, and S. Stober

Machine Learning and Knowledge Discovery in Databases

Read paper Code

Pretrained vision models often misclassify inputs that are rotated, scaled, or sheared, even though these affine transformations leave the object class unchanged. Robustness is usually restored either by building equivariance into the architecture or by retraining with augmentation, both of which require changing or retraining the model. Test-time canonicalization instead leaves the classifier untouched. It undoes the transformation of each input, mapping it to a canonical form near the training distribution before classification. Existing canonicalizers, however, rely on a narrow set of logit-based energy scores and bespoke search procedures, leaving the design space of scoring functions and optimizers unexplored. We reframe canonicalization as out-of-distribution (OOD) detection, which lets any OOD score serve as the energy minimized over transformations. Across benchmarks ranging from handwritten characters and sketches to natural images and 3D point clouds, we systematically evaluate around twenty OOD scores and nine search algorithms, finding that distance-based scores paired with random search and local refinement perform best overall. Because canonicalizing an already-aligned input can hurt accuracy, we add a gated mechanism that transforms an input only when its OOD score indicates this is needed, preserving most in-distribution accuracy while retaining the robustness gains on transformed inputs.

@inproceedings{Lindner2026,
  title     = {Zero-Shot Test-Time Canonicalization using Out-of-Distribution Scoring},
  author    = {Lindner, Dominik and Schmidt, Johann and Siegl, Tom and Becker, Martin and Stober, Sebastian},
  booktitle = {Machine Learning and Knowledge Discovery in Databases},
  doi       = {10.1007/978-3-032-37673-2_10},
  pages     = {168--185},
  year      = {2026},
  url       = {https://link.springer.com/chapter/10.1007/978-3-032-37673-2_10}
}

Speaking

Talks

Beyond writing papers, I enjoy presenting research to an audience, from conference talks and invited presentations to seminars at the university. I also took formal training in science communication and university teaching.

Community

Service

Research depends on people who read, question and guide the work of others. This is my part in it.

Coding

Stack & Tools

These are the tools I work with every day, followed by the languages I have picked up since 2014.