Source author record

Robert Kozma

Robert Kozma appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

4works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

4 published item(s)

preprint2020arXiv

Reinforcement Learning with Feedback-modulated TD-STDP

Spiking neuron networks have been used successfully to solve simple reinforcement learning tasks with continuous action set applying learning rules based on spike-timing-dependent plasticity (STDP). However, most of these models cannot be applied to reinforcement learning tasks with discrete action set since they assume that the selected action is a deterministic function of firing rate of neurons, which is continuous. In this paper, we propose a new STDP-based learning rule for spiking neuron networks which contains feedback modulation. We show that the STDP-based learning rule can be used to solve reinforcement learning tasks with discrete action set at a speed similar to standard reinforcement learning algorithms when applied to the CartPole and LunarLander tasks. Moreover, we demonstrate that the agent is unable to solve these tasks if feedback modulation is omitted from the learning rule. We conclude that feedback modulation allows better credit assignment when only the units contributing to the executed action and TD error participate in learning.

preprint2015arXiv

Bessel Functions in Mass Action. Modeling of Memories and Remembrances

Data from experimental observations of a class of neurological processes (Freeman K-sets) present functional distribution reproducing Bessel function behavior. We model such processes with couples of damped/amplified oscillators which provide time dependent representation of Bessel equation. The root loci of poles and zeros conform to solutions of K-sets. Some light is shed on the problem of filling the gap between the cellular level dynamics and the brain functional activity. Breakdown of time-reversal symmetry is related with the cortex thermodynamic features. This provides a possible mechanism to deduce lifetime of recorded memory.

preprint2015arXiv

Complete stability analysis of a heuristic ADP control design

This paper provides new stability results for Action-Dependent Heuristic Dynamic Programming (ADHDP), using a control algorithm that iteratively improves an internal model of the external world in the autonomous system based on its continuous interaction with the environment. We extend previous results by ADHDP control to the case of general multi-layer neural networks with deep learning across all layers. In particular, we show that the introduced control approach is uniformly ultimately bounded (UUB) under specific conditions on the learning rates, without explicit constraints on the temporal discount factor. We demonstrate the benefit of our results to the control of linear and nonlinear systems, including the cart-pole balancing problem. Our results show significantly improved learning and control performance as compared to the state-of-art.

preprint2012arXiv

Thermodynamic Model of Criticality in the Cortex Based On EEG/ECOG Data

Criticality in the cortex emerges from the seemingly random interaction of microscopic components and produces higher cognitive functions at mesoscopic and macroscopic scales. Random graphs and percolation theory provide natural means to de- scribe critical regions in the behavior of the cortex and they are proposed here as novel mathematical tools helping us deciphering the language of the brain.