publications

papers in reversed chronological order

2026

  1. IJCAI
    Empirical Evidence and Analysis of a Critical Pitfall in Reward Learning from Human Feedback
    Taha Shaheen ,  Stephen G. West ,  and  Yu Zhang
    In Proceedings of the 35th International Joint Conference on Artificial Intelligence (IJCAI-ECAI 2026), Aug 2026

2025

  1. arXiv
    Active Shadowing (ASD): Manipulating Visual Perception of Robotics Behaviors via Implicit Communication
    Andrew Boateng ,  Prakhar Bhartiya ,  Taha Shaheen , and 1 more author
    arXiv preprint arXiv:2407.01468, Aug 2025

2024

  1. A Survey of Reinforcement Learning from Human Feedback
    Timo Kaufmann ,  Paul Weng ,  Viktor Bengs , and 1 more author
    Transactions on Machine Learning Research, Jun 2024
  2. ACM THRI
    Investigation of Low-Moral Actions by Malicious Anonymous Operators of Avatar Robots
    Taha Shaheen ,  Dražen Brščić ,  and  Takayuki Kanda
    ACM Transactions on Human-Robot Interaction, Sep 2024

2021

  1. Robust Inverse Reinforcement Learning under Transition Dynamics Mismatch
    Luca Viano ,  Yu-Ting Huang ,  Parameswaran Kamalaruban , and 2 more authors
    In Advances in Neural Information Processing Systems, Sep 2021

2020

  1. What Is It You Really Want of Me? Generalized Reward Learning with Biased Beliefs about Domain Dynamics
    Ze Gong ,  and  Yu Zhang
    In , Apr 2020

2018

  1. Occam’s Razor Is Insufficient to Infer the Preferences of Irrational Agents
    Stuart Armstrong ,  and  Sören Mindermann
    Advances in Neural Information Processing Systems, Apr 2018

2017

  1. A Survey of Preference-Based Reinforcement Learning Methods
    Christian Wirth ,  Riad Akrour ,  Gerhard Neumann , and 1 more author
    Journal of Machine Learning Research, Apr 2017
  2. Deep Reinforcement Learning from Human Preferences
    Paul F Christiano ,  Jan Leike ,  Tom B Brown , and 3 more authors
    In , Apr 2017
  3. The off-switch game
    Dylan Hadfield-Menell ,  Anca Dragan ,  Pieter Abbeel , and 1 more author
    In Proceedings of the 26th International Joint Conference on Artificial Intelligence, Apr 2017
  4. Learning Robot Objectives from Physical Human Interaction
    Andrea Bajcsy ,  Dylan P. Losey ,  Marcia K. O’Malley , and 1 more author
    In Proceedings of the 1st Annual Conference on Robot Learning, Oct 2017

2000

  1. Algorithms for Inverse Reinforcement Learning
    Andrew Y. Ng ,  and  Stuart Russel
    In International Conference on Machine Learning, Jun 2000