Researcher profile

Rei Sato

Rei Sato contributes to research discovery and scholarly infrastructure.

ResearcherAffiliation not importedOpen to collaborate

Trust snapshot

Quick read

Trust 13 - UnverifiedVerification L1Unclaimed author
2works
0followers
3topics
4close collaborators

Actions

Decide how to stay connected

Follow researcher0

Identity and collaboration

How to connect with this researcher

Claiming links this public author record to a researcher profile and unlocks direct collaboration workflows.

Log in to claim

Direct collaboration

Open a focused conversation when the fit is right

Claim this author entity first to unlock direct invitations.

Research graph

See the researcher in context

Open full explorer

Inspect adjacent work, topics, institutions and collaborators without jumping out to a separate graph page.

Building this graph slice

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

2 published item(s)

preprint2023arXiv

Max-Min Off-Policy Actor-Critic Method Focusing on Worst-Case Robustness to Model Misspecification

In the field of reinforcement learning, because of the high cost and risk of policy training in the real world, policies are trained in a simulation environment and transferred to the corresponding real-world environment. However, the simulation environment does not perfectly mimic the real-world environment, lead to model misspecification. Multiple studies report significant deterioration of policy performance in a real-world environment. In this study, we focus on scenarios involving a simulation environment with uncertainty parameters and the set of their possible values, called the uncertainty parameter set. The aim is to optimize the worst-case performance on the uncertainty parameter set to guarantee the performance in the corresponding real-world environment. To obtain a policy for the optimization, we propose an off-policy actor-critic approach called the Max-Min Twin Delayed Deep Deterministic Policy Gradient algorithm (M2TD3), which solves a max-min optimization problem using a simultaneous gradient ascent descent approach. Experiments in multi-joint dynamics with contact (MuJoCo) environments show that the proposed method exhibited a worst-case performance superior to several baseline approaches.

preprint2020arXiv

Scaling Hypothesis of Spatial Search on Fractal Lattice Using Quantum Walk

We investigate a quantum spatial search problem on fractal lattices, such as Sierpinski carpets and Menger sponges. In earlier numerical studies of the Sierpinski gasket, the Sierpinski tetrahedron, and the Sierpinski carpet, conjectures have been proposed for the scaling of a quantum spatial search problem finding a specific target, which is given in terms of the characteristic quantities of a fractal geometry. We find that our simulation results for extended Sierpinski carpets and Menger sponges support the conjecture for the ${\it optimal}$ number of the oracle calls, where the exponent is given by $1/2$ for $d_{\rm s} > 2$ and the inverse of the spectral dimension $d_{\rm s}$ for $d_{\rm s} < 2$. We also propose a scaling hypothesis for the ${\it effective}$ number of the oracle calls defined by the ratio of the ${\it optimal}$ number of oracle calls to a square root of the maximum finding probability. The form of the scaling hypothesis for extended Sierpinski carpets is very similar but slightly different from the earlier conjecture for the Sierpinski gasket, the Sierpinski tetrahedron, and the conventional Sierpinski carpet.