Soichiro Nishimori

CV / Google Scholar / GitHub / Twitter (X)

portrait.png

About Me

I am Soichiro Nishimori, a PhD student at Sugiyama-Yokoya-Ishida lab supervised by Prof. Sugiyama. Also, I am working as a research part-timer in Imperfect-information Learning Team at RIKEN AIP.

I previously interned at OMRON SINIC X Corporation supervised by Dr. Yoshitaka Ushiku and Dr. Atsushi Hashimoto and Matsuo-Iwasawa lab at The University of Tokyo supervised by Prof. Paavo Parmas.

Mahjong 🀄️ and Tennis 🎾 lover.

Research Interests

I am interested in scalable reinforcement learning (RL) from three perspectives:

  • Scale to Data: The amount and quality of data can be traded off. I have worked on offline RL with weak supervision, including domain-unlabeled data RLC2025 and noisy preferences TMLR2026.
  • Scale to Computation: I am especially interested in GPU-accelerated RL frameworks, both for simulators, including board games ♟️ Pgx, NeurIPS2023 (Co-authored) and Riichi Mahjong 🀄️ Mahjax, arXiv2026, and for algorithm codebases, such as offline RL JAX-CORL.
  • Scale to Trial: Even if an agent can gain experience through massive data or computation, it is not truly scalable unless it can learn efficiently. From this perspective, I have worked on exploration in RL. I focus on a novel exploration objective called ReMax, based on the Retry idea (ICML2026, TMLR2026).

publications

2026

  1. Soichiro Nishimori, Shinri Okano, Keigo Habara, and 3 more authors
    arXiv preprint arXiv:2605.20577 2026
  2. Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, and 4 more authors
    International Conference on Machine Learning (ICML2026) 2026
  3. Soichiro Nishimori, Yu-Jie Zhang, Thanawat Lodkaew, and Masashi Sugiyama
    Transactions on Machine Learning Research (TMLR2026) 2026
  4. Shinnosuke Ono, Johannes Ackermann, Soichiro Nishimori, and 2 more authors
    arXiv preprint arXiv:2604.02986 2026
  5. Bingkui Tong, Junpei Komiyama, Soichiro Nishimori, and Paavo Parmas
    arXiv preprint arXiv:2605.20854 2026
  6. Shota Takashiro*, Soichiro Nishimori*, Paavo Parmas*, and 6 more authors
    arXiv preprint arXiv:2606.06080 (* Equal contribution) 2026
  7. Paavo Parmas, Yongmin Kim, Kohsei Matsutani, Shota Takashiro, Soichiro Nishimori, and 3 more authors
    arXiv preprint arXiv:2606.06096 2026
  8. Soichiro Nishimori and Paavo Parmas
    Transactions on Machine Learning Research (TMLR2026) 2026
  9. Thanawat Lodkaew, Johannes Ackermann, Soichiro Nishimori, and 3 more authors
    Advances in Neural Information Processing Systems (NeurIPS2026) 2026

2025

  1. Yuting Tang, Yivan Zhang, Johannes Ackermann, Yu-jie Zhang, Soichiro Nishimori, and 2 more authors
    Reinforcement Learning Conference (RLC2025) 2025
  2. Soichiro Nishimori, Xin-Qiang Cai, Johannes Ackermann, and Masashi Sugiyama
    Reinforcement Learning Conference (RLC2025) 2025

2024

  1. Toshinori Kitamura, Tadashi Kozuno, Masahiro Kato, Yuki Ichihara, Soichiro Nishimori, and 4 more authors
    Reinforcement Learning Conference Workshop 2024
  2. Sotetsu Koyamada, Soichiro Nishimori, and Shin Ishii
    Reinforcement Learning Conference (RLC2024) 2024
  3. Soichiro Nishimori
    github 2024

2023

  1. Sotetsu Koyamada, Shinri Okano, Soichiro Nishimori, and 4 more authors
    Advances in Neural Information Processing Systems (NeurIPS2023) 2023
  2. End-to-End Policy Gradient Method for POMDPs and Explainable Agents
    Soichiro Nishimori, Sotetsu Koyamada, and Shin Ishii
    arXiv preprint 2023

2022

  1. Sotetsu Koyamada, Keigo Habara, Nao Goto, Shinri Okano, Soichiro Nishimori, and Shin Ishii
    2022 IEEE Conference on Games (CoG) 2022

news

Sep 23, 2026 Our paper Retry Policy Gradients in Continuous Action Spaces is accepted to TMLR!
Jun 05, 2026 Preprints on Regret analysis of ReMax, ReMax in continous actions, Baseline for Max@K are out!
May 09, 2026 Our ReMax paper is accepted to ICML 2026!
Apr 27, 2026 1 paper is accepted to TMLR 2026.
Dec 26, 2025 I releaced JAX-based Mahjong simulator, Mahjax 🀄️

latest posts