{"entity": "researcher", "timestamp": "2026-08-20T20:46:58.284Z", "family": "Morita", "given": "Kenji", "initials": "K", "orcid": "0000-0003-2192-4248", "affiliations": ["Physical and Health Education, Graduate School of Education, The University of Tokyo, Tokyo, Japan.", "Theoretical Sciences Visiting Program, Okinawa Institute of Science and Technology, Okinawa, Japan.", "International Research Center for Neurointelligence (WPI-IRCN), The University of Tokyo, Tokyo, Japan."], "links": {"self": {"href": "https://publications-affiliated.scilifelab.se/researcher/5ed9416994634747b623d754cc02a5c0.json"}, "display": {"href": "https://publications-affiliated.scilifelab.se/researcher/5ed9416994634747b623d754cc02a5c0"}}, "publications": [{"entity": "publication", "iuid": "7c64d52708894e9bae5e9fdd0a485054", "links": {"self": {"href": "https://publications-affiliated.scilifelab.se/publication/7c64d52708894e9bae5e9fdd0a485054.json"}, "display": {"href": "https://publications-affiliated.scilifelab.se/publication/7c64d52708894e9bae5e9fdd0a485054"}}, "title": "Online reinforcement learning of state representation in recurrent network supported by the power of random feedback and biological constraints.", "authors": [{"family": "Tsurumi", "given": "Takayuki", "initials": "T"}, {"family": "Kato", "given": "Ayaka", "initials": "A", "orcid": "0000-0002-6306-6600", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/5b0c09585b17459a9ef2f9e40436ba81.json"}}, {"family": "Kumar", "given": "Arvind", "initials": "A", "orcid": "0000-0002-8044-9195", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/cf2fe30252074acb8d09f88f6c9b55f3.json"}}, {"family": "Morita", "given": "Kenji", "initials": "K", "orcid": "0000-0003-2192-4248", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/5ed9416994634747b623d754cc02a5c0.json"}}], "type": "journal article", "published": "2025-09-24", "journal": {"title": "Elife", "issn": "2050-084X", "volume": "14", "issn-l": "2050-084X"}, "abstract": "Representation of external and internal states in the brain plays a critical role in enabling suitable behavior. Recent studies suggest that state representation and state value can be simultaneously learned through Temporal-Difference-Reinforcement-Learning (TDRL) and Backpropagation-Through-Time (BPTT) in recurrent neural networks (RNNs) and their readout. However, neural implementation of such learning remains unclear as BPTT requires offline update using transported downstream weights, which is suggested to be biologically implausible. We demonstrate that simple online training of RNNs using TD reward prediction error and random feedback, without additional memory or eligibility trace, can still learn the structure of tasks with cue-reward delay and timing variability. This is because TD learning itself is a solution for temporal credit assignment, and feedback alignment, a mechanism originally proposed for supervised learning, enables gradient approximation without weight transport. Furthermore, we show that biologically constraining downstream weights and random feedback to be non-negative not only preserves learning but may even enhance it because the non-negative constraint ensures loose alignment-allowing the downstream and feedback weights to roughly align from the beginning. These results provide insights into the neural mechanisms underlying the learning of state representation and value, highlighting the potential of random feedback and biological constraints.", "doi": "10.7554/eLife.104101", "pmid": "40991326", "labels": [], "xrefs": [{"db": "pmc", "key": "PMC12459954"}, {"db": "pii", "key": "104101"}], "notes": [], "created": "2026-08-20T13:51:06.068Z", "modified": "2026-08-20T13:51:06.217Z"}, {"entity": "publication", "iuid": "dba3851e366e41c1b361de21a46381b3", "links": {"self": {"href": "https://publications-affiliated.scilifelab.se/publication/dba3851e366e41c1b361de21a46381b3.json"}, "display": {"href": "https://publications-affiliated.scilifelab.se/publication/dba3851e366e41c1b361de21a46381b3"}}, "title": "Online reinforcement learning of state representation in recurrent network supported by the power of random feedback and biological constraints", "authors": [{"family": "Tsurumi", "given": "Takayuki", "initials": "T"}, {"family": "Kato", "given": "Ayaka", "initials": "A", "orcid": "0000-0002-6306-6600", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/5b0c09585b17459a9ef2f9e40436ba81.json"}}, {"family": "Kumar", "given": "Arvind", "initials": "A", "orcid": "0000-0002-8044-9195", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/cf2fe30252074acb8d09f88f6c9b55f3.json"}}, {"family": "Morita", "given": "Kenji", "initials": "K", "orcid": "0000-0003-2192-4248", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/5ed9416994634747b623d754cc02a5c0.json"}}], "type": "journal-article", "published": "2025-09-24", "journal": {"issn": "2050-084X", "volume": "14", "title": "Elife", "issn-l": "2050-084X"}, "abstract": null, "doi": "10.7554/elife.104101.4", "pmid": null, "labels": [], "xrefs": [], "notes": [], "created": "2026-08-20T13:51:11.382Z", "modified": "2026-08-20T13:51:11.410Z"}]}