Source author record

Felix Schmitt

Felix Schmitt appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

10works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

10 published item(s)

preprint2022arXiv

Reward (Mis)design for Autonomous Driving

This article considers the problem of diagnosing certain common errors in reward design. Its insights are also applicable to the design of cost functions and performance metrics more generally. To diagnose common errors, we develop 8 simple sanity checks for identifying flaws in reward functions. These sanity checks are applied to reward functions from past work on reinforcement learning (RL) for autonomous driving (AD), revealing near-universal flaws in reward design for AD that might also exist pervasively across reward design for other tasks. Lastly, we explore promising directions that may aid the design of reward functions for AD in subsequent research, following a process of inquiry that can be adapted to other domains.

preprint2016arXiv

Exact Maximum Entropy Inverse Optimal Control for Modelling Human Attention Switching and Control

Maximum Causal Entropy (MCE) Inverse Optimal Control (IOC) has become an effective tool for modelling human behaviour in many control tasks. Its advantage over classic techniques for estimating human policies is the transferability of the inferred objectives: Behaviour can be predicted in variations of the control task by policy computation using a relaxed optimality criterion. However, exact policy inference is often computationally intractable in control problems with imperfect state observation. In this work, we present a model class that allows modelling human control of two tasks of which only one be perfectly observed at a time requiring attention switching. We show how efficient and exact objective and policy inference via MCE can be conducted for these control problems. Both MCE-IOC and Maximum Causal Likelihood (MCL)-IOC, a variant of the original MCE approach, as well as Direct Policy Estimation (DPE) are evaluated using simulated and real behavioural data. Prediction error and generalization over changes in the control process are both considered in the evaluation. The results show a clear advantage of both IOC methods over DPE, especially in the transfer over variation of the control process. MCE and MCL performed similar when training on a large set of simulated data, but differed significantly on small sets and real data.

preprint2016arXiv

Inverse Reinforcement Learning with Simultaneous Estimation of Rewards and Dynamics

Inverse Reinforcement Learning (IRL) describes the problem of learning an unknown reward function of a Markov Decision Process (MDP) from observed behavior of an agent. Since the agent's behavior originates in its policy and MDP policies depend on both the stochastic system dynamics as well as the reward function, the solution of the inverse problem is significantly influenced by both. Current IRL approaches assume that if the transition model is unknown, additional samples from the system's dynamics are accessible, or the observed behavior provides enough samples of the system's dynamics to solve the inverse problem accurately. These assumptions are often not satisfied. To overcome this, we present a gradient-based IRL approach that simultaneously estimates the system's dynamics. By solving the combined optimization problem, our approach takes into account the bias of the demonstrations, which stems from the generating policy. The evaluation on a synthetic MDP and a transfer learning task shows improvements regarding the sample efficiency as well as the accuracy of the estimated reward functions and transition models.

preprint2016arXiv

Predicting Lane Keeping Behavior of Visually Distracted Drivers Using Inverse Suboptimal Control

Driver distraction strongly contributes to crash-risk. Therefore, assistance systems that warn the driver if her distraction poses a hazard to road safety, promise a great safety benefit. Current approaches either seek to detect critical situations using environmental sensors or estimate a driver's attention state solely from her behavior. However, this neglects that driving situation, driver deficiencies and compensation strategies altogether determine the risk of an accident. This work proposes to use inverse suboptimal control to predict these aspects in visually distracted lane keeping. In contrast to other approaches, this allows a situation-dependent assessment of the risk posed by distraction. Real traffic data of seven drivers are used for evaluation of the predictive power of our approach. For comparison, a baseline was built using established behavior models. In the evaluation our method achieves a consistently lower prediction error over speed and track-topology variations. Additionally, our approach generalizes better to driving speeds unseen in training phase.

preprint2015arXiv

Observation of universal strong orbital-dependent correlation effects in iron chalcogenides

Establishing the appropriate theoretical framework for unconventional superconductivity in the iron-based materials requires correct understanding of both the electron correlation strength and the role of Fermi surfaces. This fundamental issue becomes especially relevant with the discovery of the iron chalcogenide (FeCh) superconductors, the only iron-based family in proximity to an insulating phase. Here, we use angle-resolved photoemission spectroscopy (ARPES) to measure three representative FeCh superconductors, FeTe0.56Se0.44, K0.76Fe1.72Se2, and monolayer FeSe film grown on SrTiO3. We show that, these FeChs are all in a strongly correlated regime at low temperatures, with an orbital-selective strong renormalization in the dxy bands despite having drastically different Fermi-surface topologies. Furthermore, raising temperature brings all three compounds from a metallic superconducting state to a phase where the dxy orbital loses all spectral weight while other orbitals remain itinerant. These observations establish that FeChs display universal orbital-selective strong correlation behaviors that are insensitive to the Fermi surface topology, and are close to an orbital-selective Mott phase (OSMP), hence placing strong constraints for theoretical understanding of iron-based superconductors.

preprint2014arXiv

Direct observation of the transition from indirect to direct bandgap in atomically thin epitaxial MoSe2

Quantum systems in confined geometries are host to novel physical phenomena. Examples include quantum Hall systems in semiconductors and Dirac electrons in graphene. Interest in such systems has also been intensified by the recent discovery of a large enhancement in photoluminescence quantum efficiency and a potential route to valleytronics in atomically thin layers of transition metal dichalcogenides, MX2 (M = Mo, W; X = S, Se, Te), which are closely related to the indirect to direct bandgap transition in monolayers. Here, we report the first direct observation of the transition from indirect to direct bandgap in monolayer samples by using angle resolved photoemission spectroscopy on high-quality thin films of MoSe2 with variable thickness, grown by molecular beam epitaxy. The band structure measured experimentally indicates a stronger tendency of monolayer MoSe2 towards a direct bandgap, as well as a larger gap size, than theoretically predicted. Moreover, our finding of a significant spin-splitting of 180 meV at the valence band maximum of a monolayer MoSe2 film could expand its possible application to spintronic devices.

preprint2013arXiv

Route-Based Detection of Conflicting ATC Clearances on Airports

Runway incursions are among the most serious safety concerns in air traffic control. Traditional A-SMGCS level 2 safety systems detect runway incursions with the help of surveillance information only. In the context of SESAR, complementary safety systems are emerging that also use other information in addition to surveillance, and that aim at warning about potential runway incursions at earlier points in time. One such system is "conflicting ATC clearances", which processes the clearances entered by the air traffic controller into an electronic flight strips system and cross-checks them for potentially dangerous inconsistencies. The cross-checking logic may be implemented directly based on the clearances and on surveillance data, but this is cumbersome. We present an approach that instead uses ground routes as an intermediate layer, thereby simplifying the core safety logic.

preprint2013arXiv

Software Design Principles of a DFS Tower A-CWP Prototype

SESAR is supposed to boost the development of new operational procedures together with the supporting systems in order to modernize the pan-European air traffic management (ATM). One consequence of this development is that more and more information is presented to - and has to be processed by - air traffic control officers (ATCOs). Thus, there is a strong need for a software design concept that fosters the development of an advanced (tower) controller working position (A-CWP) that comprehensively integrates the still counting amount of information while reducing the data management workload of ATCOs. We report on our first hands-on experiences obtained during the development of an A-CWP prototype that was used in two SESAR validation sessions.

preprint2010arXiv

Ab-initio phase diagram of ultracold 87-Rb in an one-dimensional two-color superlattice

We investigate the ab-initio phase diagram of ultracold 87-Rb atoms in an one-dimensional two-color superlattice. Using single-particle band structure calculations we map the experimental setup onto the parameters of the Bose-Hubbard model. This ab-initio ansatz allows us to express the phase diagrams in terms of the experimental control parameters, i.e., the intensities of the lasers that form the optical superlattice. In order to solve the many-body problem for experimental system sizes we adopt the density-matrix renormalization-group algorithm. A detailed study of convergence and finite-size effects for all observables is presented. Our results show that all relevant quantum phases, i.e., superfluid, Mott-insulator, and quasi Bose-glass, can be accessed through intensity variation of the lasers alone. However, it turns out that the phase diagram is strongly affected by the longitudinal trapping potential.