Source author record

Youngchul Sung

Youngchul Sung appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

26works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

26 published item(s)

preprint2026arXiv

Adaptive Action Chunking via Multi-Chunk Q Value Estimation

Action chunking emerged as a pivotal technique in imitation learning, enabling policies to predict cohesive action sequences rather than single actions. Recently, this approach has expanded to reinforcement learning (RL), enhancing behavioral consistency and reducing bootstrapping errors in value function estimation. However, existing methods rely on a fixed chunk length, creating a performance bottleneck as the optimal length varies across states and tasks. In this paper, we propose Adaptive Action CHunking (ACH), a novel offline-to-online RL algorithm that dynamically modulates chunk length during both training and inference. To find the optimal chunk length for a dynamically varying current state, we simultaneously estimate action-values for all candidate chunk lengths in a single forward pass, using a Transformer-based architecture. Our mechanism allows the agent to select the most effective chunk length adaptively based on the current state. Evaluated on 34 challenging tasks, ACH consistently outperforms fixed-length baselines, demonstrating superior generalization and learning efficiency in complex environments.

preprint2025arXiv

Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied Agents

In this paper, we propose a test-time adaptive agent that performs exploratory inference through posterior-guided belief refinement without relying on gradient-based updates or additional training for LLM agent operating under partial observability. Our agent maintains an external structured belief over the environment state, iteratively updates it via action-conditioned observations, and selects actions by maximizing predicted information gain over the belief space. We estimate information gain using a lightweight LLM-based surrogate and assess world alignment through a novel reward that quantifies the consistency between posterior belief and ground-truth environment configuration. Experiments show that our method outperforms inference-time scaling baselines such as prompt-augmented or retrieval-enhanced LLMs, in aligning with latent world states with significantly lower integration overhead.

preprint2022arXiv

MASER: Multi-Agent Reinforcement Learning with Subgoals Generated from Experience Replay Buffer

In this paper, we consider cooperative multi-agent reinforcement learning (MARL) with sparse reward. To tackle this problem, we propose a novel method named MASER: MARL with subgoals generated from experience replay buffer. Under the widely-used assumption of centralized training with decentralized execution and consistent Q-value decomposition for MARL, MASER automatically generates proper subgoals for multiple agents from the experience replay buffer by considering both individual Q-value and total Q-value. Then, MASER designs individual intrinsic reward for each agent based on actionable representation relevant to Q-learning so that the agents reach their subgoals while maximizing the joint action value. Numerical results show that MASER significantly outperforms StarCraft II micromanagement benchmark compared to other state-of-the-art MARL algorithms.

preprint2022arXiv

Robust Imitation Learning against Variations in Environment Dynamics

In this paper, we propose a robust imitation learning (IL) framework that improves the robustness of IL when environment dynamics are perturbed. The existing IL framework trained in a single environment can catastrophically fail with perturbations in environment dynamics because it does not capture the situation that underlying environment dynamics can be changed. Our framework effectively deals with environments with varying dynamics by imitating multiple experts in sampled environment dynamics to enhance the robustness in general variations in environment dynamics. In order to robustly imitate the multiple sample experts, we minimize the risk with respect to the Jensen-Shannon divergence between the agent's policy and each of the sample experts. Numerical results show that our algorithm significantly improves robustness against dynamics perturbations compared to conventional IL baselines.

preprint2020arXiv

A Maximum Mutual Information Framework for Multi-Agent Reinforcement Learning

In this paper, we propose a maximum mutual information (MMI) framework for multi-agent reinforcement learning (MARL) to enable multiple agents to learn coordinated behaviors by regularizing the accumulated return with the mutual information between actions. By introducing a latent variable to induce nonzero mutual information between actions and applying a variational bound, we derive a tractable lower bound on the considered MMI-regularized objective function. Applying policy iteration to maximize the derived lower bound, we propose a practical algorithm named variational maximum mutual information multi-agent actor-critic (VM3-AC), which follows centralized learning with decentralized execution (CTDE). We evaluated VM3-AC for several games requiring coordination, and numerical results show that VM3-AC outperforms MADDPG and other MARL algorithms in multi-agent tasks requiring coordination.

preprint2020arXiv

Population-Guided Parallel Policy Search for Reinforcement Learning

In this paper, a new population-guided parallel learning scheme is proposed to enhance the performance of off-policy reinforcement learning (RL). In the proposed scheme, multiple identical learners with their own value-functions and policies share a common experience replay buffer, and search a good policy in collaboration with the guidance of the best policy information. The key point is that the information of the best policy is fused in a soft manner by constructing an augmented loss function for policy update to enlarge the overall search region by the multiple learners. The guidance by the previous best policy and the enlarged range enable faster and better policy search. Monotone improvement of the expected cumulative return by the proposed scheme is proved theoretically. Working algorithms are constructed by applying the proposed scheme to the twin delayed deep deterministic (TD3) policy gradient algorithm. Numerical results show that the constructed algorithm outperforms most of the current state-of-the-art RL algorithms, and the gain is significant in the case of sparse reward environment.

preprint2015arXiv

Enhancing Non-Orthogonal Multiple Access By Forming Relaying Broadcast Channels

In this paper, using relaying broadcast channels (RBCs) as component channels for non-orthogonal multiple access (NOMA) is proposed to enhance the performance of NOMA in single-input single-output (SISO) cellular downlink systems. To analyze the performance of the proposed scheme, an achievable rate region of a RBC with compress-and-forward (CF) relaying is newly derived based on the recent work of noisy network coding (NNC). Based on the analysis of the achievable rate region of a RBC with decode-and-forward (DF) relaying, CF relaying, or CF relaying with dirty-paper coding (DPC) at the transmitter, the overall system performance of NOMA equipped with RBC component channels is investigated. It is shown that NOMA with RBC-DF yields marginal gain and NOMA with RBC-CF/DPC yields drastic gain over the simple NOMA based on broadcast component channels in a practical system setup. By going beyond simple broadcast channel (BC)/successive interference cancellation (SIC) to advanced multi-terminal encoding including DPC and CF/NNC, far larger gains can be obtained for NOMA.

preprint2015arXiv

Randomly-Directional Beamforming in Millimeter-Wave Multi-User MISO Downlink

In this paper, randomly-directional beamforming (RDB) is considered for millimeter-wave (mmwave) multi-user (MU) multiple-input single-output (MISO) downlink systems. By using asymptotic techniques, the performance of RDB and the MU gain in mm-wave MISO are analyzed based on the uniform random line-of-sight (UR-LoS) channel model suitable for highly directional mm-wave radio propagation channels. It is shown that there exists a transition point on the number of users relative to the number of antenna elements for non-trivial performance of the RDB scheme, and furthermore sum rate scaling arbitrarily close to linear scaling with respect to the number of antenna elements can be achieved under the UR-LoS channel model by opportunistic random beamforming with proper user scheduling if the number of users increases linearly with respect to the number of antenna elements. The provided results yield insights into the most effective beamforming and scheduling choices for mm-wave MU-MISO in various operating conditions. Simulation results validate our analysis based on asymptotic techniques for finite cases.

preprint2014arXiv

A New Approach to User Scheduling in Massive Multi-User MIMO Broadcast Channels

In this paper, a new user-scheduling-and-beamforming method is proposed for multi-user massive multiple-input multiple-output (massive MIMO) broadcast channels in the context of two-stage beamforming. The key ideas of the proposed scheduling method are 1) to use a set of orthogonal reference beams and construct a double cone around each reference beam to select `nearly-optimal' semi-orthogonal users based only on channel quality indicator (CQI) feedback and 2) to apply post-user-selection beam refinement with zero-forcing beamforming (ZFBF) based on channel state information (CSI) feedback only from the selected users. It is proved that the proposed scheduling-and-beamforming method is asymptotically optimal as the number of users increases. Furthermore, the proposed scheduling-and-beamforming method almost achieves the performance of the existing semi-orthogonal user selection with ZFBF (SUS-ZFBF) that requires full CSI feedback from all users, with significantly reduced feedback overhead which is even less than that required by random beamforming.

preprint2014arXiv

Pilot Beam Pattern Design for Channel Estimation in Massive MIMO Systems

In this paper, the problem of pilot beam pattern design for channel estimation in massive multiple-input multiple-output systems with a large number of transmit antennas at the base station is considered, and a new algorithm for pilot beam pattern design for optimal channel estimation is proposed under the assumption that the channel is a stationary Gauss-Markov random process. The proposed algorithm designs the pilot beam pattern sequentially by exploiting the properties of Kalman filtering and the associated prediction error covariance matrices and also the channel statistics such as spatial and temporal channel correlation. The resulting design generates a sequentially-optimal sequence of pilot beam patterns with low complexity for a given set of system parameters. Numerical results show the effectiveness of the proposed algorithm.

preprint2014arXiv

Pilot Beam Sequence Design for Channel Estimation in Millimeter-Wave MIMO Systems: A POMDP Framework

In this paper, adaptive pilot beam sequence design for channel estimation in large millimeter-wave (mmWave) MIMO systems is considered. By exploiting the sparsity of mmWave MIMO channels with the virtual channel representation and imposing a Markovian random walk assumption on the physical movement of the line-of-sight (LOS) and reflection clusters, it is shown that the sparse channel estimation problem in large mmWave MIMO systems reduces to a sequential detection problem that finds the locations and values of the non-zero-valued bins in a two-dimensional rectangular grid, and the optimal adaptive pilot design problem can be cast into the framework of a partially observable Markov decision process (POMDP). Under the POMDP framework, an optimal adaptive pilot beam sequence design method is obtained to maximize the accumulated transmission data rate for a given period of time. Numerical results are provided to validate our pilot signal design method and they show that the proposed method yields good performance.

preprint2014arXiv

Pilot Signal Design for Massive MIMO Systems: A Received Signal-To-Noise-Ratio-Based Approach

In this paper, the pilot signal design for massive MIMO systems to maximize the training-based received signal-to-noise ratio (SNR) is considered under two channel models: block Gauss-Markov and block independent and identically distributed (i.i.d.) channel models. First, it is shown that under the block Gauss-Markov channel model, the optimal pilot design problem reduces to a semi-definite programming (SDP) problem, which can be solved numerically by a standard convex optimization tool. Second, under the block i.i.d. channel model, an optimal solution is obtained in closed form. Numerical results show that the proposed method yields noticeably better performance than other existing pilot design methods in terms of received SNR.

preprint2014arXiv

Training Beam Sequence Design for Millimeter-Wave MIMO Systems: A POMDP Framework

In this paper, adaptive training beam sequence design for efficient channel estimation in large millimeter-wave(mmWave) multiple-input multiple-output (MIMO) channels is considered. By exploiting the sparsity in large mmWave MIMO channels and imposing a Markovian random walk assumption on the movement of the receiver and reflection clusters, the adaptive training beam sequence design and channel estimation problem is formulated as a partially observableMarkov decision process (POMDP) problem that finds non-zero bins in a two-dimensional grid. Under the proposed POMDP framework, optimal and suboptimal adaptive training beam sequence design policies are derived. Furthermore, a very fast suboptimal greedy algorithm is developed based on a newly proposed reduced sufficient statistic to make the computational complexity of the proposed algorithm low to a level for practical implementation. Numerical results are provided to evaluate the performance of the proposed training beam design method. Numerical results show that the proposed training beam sequence design algorithms yield good performance.

preprint2014arXiv

Two-Stage Beamformer Design for Massive MIMO Downlink By Trace Quotient Formulation

In this paper, the problem of outer beamformer design based only on channel statistic information is considered for two-stage beamforming for multi-user massive MIMO downlink, and the problem is approached based on signal-to-leakage-plus-noise ratio (SLNR). To eliminate the dependence on the instantaneous channel state information, a lower bound on the average SLNR is derived by assuming zero-forcing (ZF) inner beamforming, and an outer beamformer design method that maximizes the lower bound on the average SLNR is proposed. It is shown that the proposed SLNR-based outer beamformer design problem reduces to a trace quotient problem (TQP), which is often encountered in the field of machine learning. An iterative algorithm is presented to obtain an optimal solution to the proposed TQP. The proposed method has the capability of optimally controlling the weighting factor between the signal power to the desired user and the interference leakage power to undesired users according to different channel statistics. Numerical results show that the proposed outer beamformer design method yields significant performance gain over existing methods.

preprint2013arXiv

Filter-And-Forward Relay Design for MIMO-OFDM Systems

In this paper, the filter-and-forward (FF) relay design for multiple-input multiple-output (MIMO) orthogonal frequency-division multiplexing (OFDM) systems is considered. Due to the considered MIMO structure, the problem of joint design of the linear MIMO transceiver at the source and the destination and the FF relay at the relay is considered. As the design criterion, the minimization of weighted sum mean-square-error (MSE) is considered first, and the joint design in this case is approached based on alternating optimization that iterates between optimal design of the FF relay for a given set of MIMO precoder and decoder and optimal design of the MIMO precoder and decoder for a given FF relay filter. Next, the joint design problem for rate maximization is considered based on the obtained result regarding weighted sum MSE and the existing result regarding the relationship between weighted MSE minimization and rate maximization. Numerical results show the effectiveness of the proposed FF relay design and significant performance improvement by FF relays over widely-considered simple AF relays for MIMO-ODFM systems.

preprint2013arXiv

Filter-and-Forward Transparent Relay Design for OFDM Systems

In this paper, the filter-and-forward (FF) relay design for orthogonal frequency-division multiplexing (OFDM) transmission systems is considered to improve the system performance over simple amplify-and-forward (AF) relaying. Unlike conventional OFDM relays performing OFDM demodulation and remodulation, to reduce processing complexity, the proposed FF relay directly filters the incoming signal in time domain with a finite impulse response (FIR) and forwards the filtered signal to the destination. Three design criteria are considered to optimize the relay filter. The first criterion is the minimization of the relay transmit power subject to per-subcarrier signal-to-noise ratio (SNR) constraints, the second is the maximization of the worst subcarrier channel SNR subject to source and relay transmit power constraints, and the third is the maximization of data rate subject to source and relay transmit power constraints. It is shown that the first problem reduces to a semi-definite programming (SDP) problem by semi-definite relaxation and the solution to the relaxed SDP problem has rank one under a mild condition. For the latter two problems, the problem of joint source power allocation and relay filter design is considered and an efficient algorithm is proposed for each problem based on alternating optimization and the projected gradient method (PGM). Numerical results show that the proposed FF relay significantly outperforms simple AF relays with insignificant increase in complexity. Thus, the proposed FF relay provides a practical alternative to the AF relaying scheme for OFDM transmission.

preprint2012arXiv

Coordinated Beamforming with Relaxed Zero Forcing: The Sequential Orthogonal Projection Combining Method and Rate Control

In this paper, coordinated beamforming based on relaxed zero forcing (RZF) for K transmitter-receiver pair multiple-input single-output (MISO) and multiple-input multiple-output (MIMO) interference channels is considered. In the RZF coordinated beamforming, conventional zero-forcing interference leakage constraints are relaxed so that some predetermined interference leakage to undesired receivers is allowed in order to increase the beam design space for larger rates than those of the zero-forcing (ZF) scheme or to make beam design feasible when ZF is impossible. In the MISO case, it is shown that the rate-maximizing beam vector under the RZF framework for a given set of interference leakage levels can be obtained by sequential orthogonal projection combining (SOPC). Based on this, exact and approximate closed-form solutions are provided in two-user and three-user cases, respectively, and an efficient beam design algorithm for RZF coordinated beamforming is provided in general cases. Furthermore, the rate control problem under the RZF framework is considered. A centralized approach and a distributed heuristic approach are proposed to control the position of the designed rate-tuple in the achievable rate region. Finally, the RZF framework is extended to MIMO interference channels by deriving a new lower bound on the rate of each user.

preprint2012arXiv

On the Pareto-Optimal Beam Structure and Design for Multi-User MIMO Interference Channels

In this paper, the Pareto-optimal beam structure for multi-user multiple-input multiple-output (MIMO) interference channels is investigated and a necessary condition for any Pareto-optimal transmit signal covariance matrix is presented for the K-pair Gaussian (N,M_1,...,M_K) interference channel. It is shown that any Pareto-optimal transmit signal covariance matrix at a transmitter should have its column space contained in the union of the eigen-spaces of the channel matrices from the transmitter to all receivers. Based on this necessary condition, an efficient parameterization for the beam search space is proposed. The proposed parameterization is given by the product manifold of a Stiefel manifold and a subset of a hyperplane and enables us to construct a very efficient beam design algorithm by exploiting its rich geometrical structure and existing tools for optimization on Stiefel manifolds. Reduction in the beam search space dimension and computational complexity by the proposed parameterization and the proposed beam design approach is significant when the number of transmit antennas is larger than the sum of the numbers of receive antennas, as in upcoming cellular networks adopting massive MIMO technologies. Numerical results validate the proposed parameterization and the proposed cooperative beam design method based on the parameterization for MIMO interference channels.

preprint2012arXiv

Outage Probability and Outage-Based Robust Beamforming for MIMO Interference Channels with Imperfect Channel State Information

In this paper, the outage probability and outage-based beam design for multiple-input multiple-output (MIMO) interference channels are considered. First, closed-form expressions for the outage probability in MIMO interference channels are derived under the assumption of Gaussian-distributed channel state information (CSI) error, and the asymptotic behavior of the outage probability as a function of several system parameters is examined by using the Chernoff bound. It is shown that the outage probability decreases exponentially with respect to the quality of CSI measured by the inverse of the mean square error of CSI. Second, based on the derived outage probability expressions, an iterative beam design algorithm for maximizing the sum outage rate is proposed. Numerical results show that the proposed beam design algorithm yields better sum outage rate performance than conventional algorithms such as interference alignment developed under the assumption of perfect CSI.

preprint2011arXiv

A joint time-invariant filtering approach to the linear Gaussian relay problem

In this paper, the linear Gaussian relay problem is considered. Under the linear time-invariant (LTI) model the problem is formulated in the frequency domain based on the Toeplitz distribution theorem. Under the further assumption of realizable input spectra, the LTI Gaussian relay problem is converted to a joint design problem of source and relay filters under two power constraints, one at the source and the other at the relay, and a practical solution to this problem is proposed based on the projected subgradient method. Numerical results show that the proposed method yields a noticeable gain over the instantaneous amplify-and-forward (AF) scheme in inter-symbol interference (ISI) channels. Also, the optimality of the AF scheme within the class of one-tap relay filters is established in flat-fading channels.

preprint2011arXiv

The capacity for the linear time-invariant Gaussian relay channel

In this paper, the Gaussian relay channel with linear time-invariant relay filtering is considered. Based on spectral theory for stationary processes, the maximum achievable rate for this subclass of linear Gaussian relay operation is obtained in finite-letter characterization. The maximum rate can be achieved by dividing the overall frequency band into at most eight subbands and by making the relay behave as an instantaneous amplify-and-forward relay at each subband. Numerical results are provided to evaluate the performance of LTI relaying.

preprint2010arXiv

On Beamformer Design for Multiuser MIMO Interference Channels

This paper considers several linear beamformer design paradigms for multiuser time-invariant multiple-input multiple-output interference channels. Notably, interference alignment and sum-rate based algorithms such as the maximum signal-to-interference-plus noise (max-SINR) algorithm are considered. Optimal linear beamforming under interference alignment consists of two layers; an inner precoder and decoder (or receive filter) accomplish interference alignment to eliminate inter-user interference, and an outer precoder and decoder diagonalize the effective single-user channel resulting from the interference alignment by the inner precoder and decoder. The relationship between this two-layer beamforming and the max-SINR algorithm is established at high signal-to-noise ratio. Also, the optimality of the max-SINR algorithm within the class of linear beamforming algorithms, and its local convergence with exponential rate, are established at high signal-to-noise ratio.

preprint2008arXiv

Information, Energy and Density for Ad Hoc Sensor Networks over Correlated Random Fields: Large Deviations Analysis

Using large deviations results that characterize the amount of information per node on a two-dimensional (2-D) lattice, asymptotic behavior of a sensor network deployed over a correlated random field for statistical inference is investigated. Under a 2-D hidden Gauss-Markov random field model with symmetric first order conditional autoregression, the behavior of the total information [nats] and energy efficiency [nats/J] defined as the ratio of total gathered information to the required energy is obtained as the coverage area, node density and energy vary.

preprint2008arXiv

Large Deviations Analysis for the Detection of 2D Hidden Gauss-Markov Random Fields Using Sensor Networks

The detection of hidden two-dimensional Gauss-Markov random fields using sensor networks is considered. Under a conditional autoregressive model, the error exponent for the Neyman-Pearson detector satisfying a fixed level constraint is obtained using the large deviations principle. For a symmetric first order autoregressive model, the error exponent is given explicitly in terms of the SNR and an edge dependence factor (field correlation). The behavior of the error exponent as a function of correlation strength is seen to divide into two regions depending on the value of the SNR. At high SNR, uncorrelated observations maximize the error exponent for a given SNR, whereas there is non-zero optimal correlation at low SNR. Based on the error exponent, the energy efficiency (defined as the ratio of the total information gathered to the total energy required) of ad hoc sensor network for detection is examined for two sensor deployment models: an infinite area model and and infinite density model. For a fixed sensor density, the energy efficiency diminishes to zero at rate O(area^{-1/2}) as the area is increased. On the other hand, non-zero efficiency is possible for increasing density depending on the behavior of the physical correlation as a function of the link length.

preprint2008arXiv

Optimal Node Density for Two-Dimensional Sensor Arrays

The problem of optimal node density for ad hoc sensor networks deployed for making inferences about two dimensional correlated random fields is considered. Using a symmetric first order conditional autoregressive Gauss-Markov random field model, large deviations results are used to characterize the asymptotic per-node information gained from the array. This result then allows an analysis of the node density that maximizes the information under an energy constraint, yielding insights into the trade-offs among the information, density and energy.

preprint2006arXiv

Neyman-Pearson Detection of Gauss-Markov Signals in Noise: Closed-Form Error Exponent and Properties

The performance of Neyman-Pearson detection of correlated stochastic signals using noisy observations is investigated via the error exponent for the miss probability with a fixed level. Using the state-space structure of the signal and observation model, a closed-form expression for the error exponent is derived, and the connection between the asymptotic behavior of the optimal detector and that of the Kalman filter is established. The properties of the error exponent are investigated for the scalar case. It is shown that the error exponent has distinct characteristics with respect to correlation strength: for signal-to-noise ratio (SNR) >1 the error exponent decreases monotonically as the correlation becomes stronger, whereas for SNR <1 there is an optimal correlation that maximizes the error exponent for a given SNR.