# RL4J

**URL:** https://community.konduit.ai/c/rl4j/12.md

[Latest](https://community.konduit.ai/latest.md) · [Categories](https://community.konduit.ai/categories.md)

---

## [About the RL4J category](https://community.konduit.ai/t/about-the-rl4j-category/36)

<div class="topic-metadata">

**Author:** [@treo](https://community.konduit.ai/u/treo)\
**Replies:** 0

</div>

RL4J: Reinforcement Learning for Java RL4J is still a tech preview. So expect to find rough edges and not quite great documentation.

---

## [Is there any example for doing policy gradient calculation with dl4j/rl4j?](https://community.konduit.ai/t/is-there-any-example-for-doing-policy-gradient-calculation-with-dl4j-rl4j/240)

<div class="topic-metadata">

**Author:** [@RuralHunter](https://community.konduit.ai/u/RuralHunter)\
**Replies:** 14\
**Last updated:** [December 11, 2024, 6:34am UTC](https://community.konduit.ai/t/is-there-any-example-for-doing-policy-gradient-calculation-with-dl4j-rl4j/240 "2024-12-11T06:34:06Z")

</div>

I can not find any. Since dl4j embeds the activation function in the layer, it seems difficult to calculate the gradient externally and put it back in dl4j network.

---

## [Is there a way to load a QLearningDiscreteDense Object after saving it?](https://community.konduit.ai/t/is-there-a-way-to-load-a-qlearningdiscretedense-object-after-saving-it/2835)

<div class="topic-metadata">

**Author:** [@Alban](https://community.konduit.ai/u/Alban)\
**Replies:** 0\
**Last updated:** [October 3, 2023, 1:11am UTC](https://community.konduit.ai/t/is-there-a-way-to-load-a-qlearningdiscretedense-object-after-saving-it/2835 "2023-10-03T01:11:13Z")

</div>

I trained a QLearningDiscreteDense object on a set of data and saved it. After calling .getNeuralNet().save(networkName) on that object I had a zip file for my Network. Is there a way to load that network again and train…

---

## [DenseLayer (index=9, name=ffn0) nIn=0, nOut=3072; nIn and nOut must be \> 0](https://community.konduit.ai/t/denselayer-index-9-name-ffn0-nin-0-nout-3072-nin-and-nout-must-be-0/2707)

<div class="topic-metadata">

**Author:** [@whaile-off](https://community.konduit.ai/u/whaile-off)\
**Replies:** 18\
**Last updated:** [August 1, 2023, 11:54pm UTC](https://community.konduit.ai/t/denselayer-index-9-name-ffn0-nin-0-nout-3072-nin-and-nout-must-be-0/2707 "2023-08-01T23:54:31Z")

</div>

Help I have error in my code public class Main { private static final Logger log = LoggerFactory.getLogger(Main.class); private static final int batchSize = 15; public static void main(String\[\] args) throws Exception…

---

## [Where are the RL4J examples?](https://community.konduit.ai/t/where-are-the-rl4j-examples/1864)

<div class="topic-metadata">

**Author:** [@GPSforLEGENDS](https://community.konduit.ai/u/GPSforLEGENDS)\
**Replies:** 4\
**Last updated:** [January 5, 2023, 2:16am UTC](https://community.konduit.ai/t/where-are-the-rl4j-examples/1864 "2023-01-05T02:16:21Z")

</div>

Hi, does anyone know where the examples for RL4J are? And is there any documentation for RL4J?

---

## [Follow the instructions to build RL4J but error reported](https://community.konduit.ai/t/follow-the-instructions-to-build-rl4j-but-error-reported/2043)

<div class="topic-metadata">

**Author:** [@TempKonduitUser1](https://community.konduit.ai/u/TempKonduitUser1)\
**Replies:** 3\
**Last updated:** [October 17, 2022, 12:02pm UTC](https://community.konduit.ai/t/follow-the-instructions-to-build-rl4j-but-error-reported/2043 "2022-10-17T12:02:18Z")

</div>

Have followed the instructions to build RL4J from RL4J by the instruction: mvn install but it reported the following error: … \[INFO\] ------------------------------------------------------------------------ \[INFO\] Re…

---

## [Failed to build RL4J with erros](https://community.konduit.ai/t/failed-to-build-rl4j-with-erros/2041)

<div class="topic-metadata">

**Author:** [@TempKonduitUser1](https://community.konduit.ai/u/TempKonduitUser1)\
**Replies:** 0\
**Last updated:** [October 14, 2022, 2:54pm UTC](https://community.konduit.ai/t/failed-to-build-rl4j-with-erros/2041 "2022-10-14T14:54:29Z")

</div>

Have followed the direction from here: \[GitHub - rubenfiszel/rl4j: Deep Reinforcement Learning for the JVM (Deep-Q, A3C)\](https://rubenfisze’s RL4J), by run the following command: mvn install -pl rl4j-api in the github…

---

## [Problem removing lstm, shape exception](https://community.konduit.ai/t/problem-removing-lstm-shape-exception/1918)

<div class="topic-metadata">

**Author:** [@petter](https://community.konduit.ai/u/petter)\
**Replies:** 5\
**Last updated:** [July 3, 2022, 11:52pm UTC](https://community.konduit.ai/t/problem-removing-lstm-shape-exception/1918 "2022-07-03T23:52:56Z")

</div>

Problem removing lstm, shape exception I am running some tests using rl4j with A3C applied to a gym-environment in OpenAI. I am running on a Windows machine and using latest versions (1.0.0-M1.1) of dl4j and rl4j. The…

---

## [Rl4j moved out of dl4j, what is the status?](https://community.konduit.ai/t/rl4j-moved-out-of-dl4j-what-is-the-status/1914)

<div class="topic-metadata">

**Author:** [@petter](https://community.konduit.ai/u/petter)\
**Replies:** 2\
**Last updated:** [July 2, 2022, 5:22pm UTC](https://community.konduit.ai/t/rl4j-moved-out-of-dl4j-what-is-the-status/1914 "2022-07-02T17:22:12Z")

</div>

I was (more or less happily) using version 1.0.0-M1.1 of dl4j and decided to “upgrade” to latest version: 1.0.0-M2 and soon realized that the entire rl4j, which I am relying on, was moved out of dl4j. So now to my questi…

---

## [Serialization with RL4j](https://community.konduit.ai/t/serialization-with-rl4j/1899)

<div class="topic-metadata">

**Author:** [@marcus.frex](https://community.konduit.ai/u/marcus.frex)\
**Replies:** 3\
**Last updated:** [June 23, 2022, 7:40am UTC](https://community.konduit.ai/t/serialization-with-rl4j/1899 "2022-06-23T07:40:47Z")

</div>

Hello everyone, I am working with NStepQLearning model but it seems just saving ComputationGraph it is using is not enough because when i restored the model it does not give the same results given before. Do you guys h…

---

## [Migrating RL4J to DL4J Contrib](https://community.konduit.ai/t/migrating-rl4j-to-dl4j-contrib/1792)

<div class="topic-metadata">

**Author:** [@kgoderis](https://community.konduit.ai/u/kgoderis)\
**Replies:** 2\
**Last updated:** [March 3, 2022, 7:45am UTC](https://community.konduit.ai/t/migrating-rl4j-to-dl4j-contrib/1792 "2022-03-03T07:45:17Z")

</div>

Hi all I have been working on updating RL4J since a few weeks, and things are getting in a state that it can be reviewed, and benefit from your feedback. In short, this is what has been done Generalise the generic cla…

---

## [Trying to use QLearning in a custom MDP environment. Chooses action 0 every time, despite the heavy negative reward](https://community.konduit.ai/t/trying-to-use-qlearning-in-a-custom-mdp-environment-chooses-action-0-every-time-despite-the-heavy-negative-reward/1024)

<div class="topic-metadata">

**Author:** [@jonathonbird810](https://community.konduit.ai/u/jonathonbird810)\
**Replies:** 22\
**Last updated:** [May 6, 2021, 6:22pm UTC](https://community.konduit.ai/t/trying-to-use-qlearning-in-a-custom-mdp-environment-chooses-action-0-every-time-despite-the-heavy-negative-reward/1024 "2021-05-06T18:22:21Z")

</div>

It’s playing a dots and boxes game. I have been messing around with the hyperparameters constantly but nothing changes. I’m completely stuck and need help. If anyone could look through the code and see if anything is set…

---

## [RL4L Documentation Request](https://community.konduit.ai/t/rl4l-documentation-request/1319)

<div class="topic-metadata">

**Author:** [@ssb](https://community.konduit.ai/u/ssb)\
**Replies:** 0\
**Last updated:** [April 15, 2021, 11:54pm UTC](https://community.konduit.ai/t/rl4l-documentation-request/1319 "2021-04-15T23:54:34Z")

</div>

I have found the RL4J examples, but would like more explanation about the toy game. I am creating my own MDP environment (not using a predefined gym game or MALMO). Is there documentation that explains the methods and …

---

## [OutOfMemoryError and Memory Management](https://community.konduit.ai/t/outofmemoryerror-and-memory-management/1274)

<div class="topic-metadata">

**Author:** [@berse2212](https://community.konduit.ai/u/berse2212)\
**Replies:** 0\
**Last updated:** [March 24, 2021, 4:13pm UTC](https://community.konduit.ai/t/outofmemoryerror-and-memory-management/1274 "2021-03-24T16:13:55Z")

</div>

Hi, I am currently working on a scenario with a somewhat bigger observation space (541 to be exact). Unfortunately this leads to me running into OutOfMemoryErrors at roughly 16G. This is basically half of my RAM. I rea…

---

## [Encodable and ObservationSpace](https://community.konduit.ai/t/encodable-and-observationspace/1125)

<div class="topic-metadata">

**Author:** [@berse2212](https://community.konduit.ai/u/berse2212)\
**Replies:** 5\
**Last updated:** [January 17, 2021, 8:40pm UTC](https://community.konduit.ai/t/encodable-and-observationspace/1125 "2021-01-17T20:40:40Z")

</div>

Hi, I am fairly new to rl4j and reinforcement learning in general. I tried to implement a very simple MDP (Snake) to gain more experience. However since the documentation is still WOP I have a problem comprehending how …

---

## [My Network isn't Learning anything and A3C keeps crashing](https://community.konduit.ai/t/my-network-isnt-learning-anything-and-a3c-keeps-crashing/326)

<div class="topic-metadata">

**Author:** [@JavaLearningYT](https://community.konduit.ai/u/JavaLearningYT)\
**Replies:** 5\
**Last updated:** [December 20, 2020, 6:38pm UTC](https://community.konduit.ai/t/my-network-isnt-learning-anything-and-a3c-keeps-crashing/326 "2020-12-20T18:38:09Z")

</div>

I am trying to teach a network how to play snake. I am using 1.0.0-beta6 I have used QLearning Async QLearning Actor-Critic None of these systems have had any luck. After millions of steps on Async QLearning…

---

## [A3C training on asynchronous game](https://community.konduit.ai/t/a3c-training-on-asynchronous-game/820)

<div class="topic-metadata">

**Author:** [@CountZukula](https://community.konduit.ai/u/CountZukula)\
**Replies:** 4\
**Last updated:** [August 28, 2020, 7:38am UTC](https://community.konduit.ai/t/a3c-training-on-asynchronous-game/820 "2020-08-28T07:38:20Z")

</div>

Hi all, I’m getting to grips with the concept of neural nets and would like to apply RL4J to a game we have running in our lab. The 2D game is grid based, with a server calling clients/players asynchronously through a me…

---

## [Forked Gym server Repo](https://community.konduit.ai/t/forked-gym-server-repo/800)

<div class="topic-metadata">

**Author:** [@cagneymoreau](https://community.konduit.ai/u/cagneymoreau)\
**Replies:** 2\
**Last updated:** [August 24, 2020, 10:10pm UTC](https://community.konduit.ai/t/forked-gym-server-repo/800 "2020-08-24T22:10:21Z")

</div>

If anyone has suffered through connecting the gym to dl4j this might help them. Its a python 3.\* and tensorflow 2.\* compatable gym server

---

## [Simplified example](https://community.konduit.ai/t/simplified-example/621)

<div class="topic-metadata">

**Author:** [@sigmund](https://community.konduit.ai/u/sigmund)\
**Replies:** 2\
**Last updated:** [June 16, 2020, 11:19am UTC](https://community.konduit.ai/t/simplified-example/621 "2020-06-16T11:19:49Z")

</div>

Hi, i try to create a simplified rl4j example based on the existing Gym and Malmo examples. Given is a sine wave and the AI should say if we are on top of the wave, on bottom or somewhere else(noop). The SineRider is t…

---

## [Issue with nextAction(Observation observation) on beta7](https://community.konduit.ai/t/issue-with-nextaction-observation-observation-on-beta7/577)

<div class="topic-metadata">

**Author:** [@GaboGoyx](https://community.konduit.ai/u/GaboGoyx)\
**Replies:** 2\
**Last updated:** [June 3, 2020, 2:11pm UTC](https://community.konduit.ai/t/issue-with-nextaction-observation-observation-on-beta7/577 "2020-06-03T14:11:12Z")

</div>

Hi, I’m having trouble with nextAction(…) from DQNPolicy on 1.0.0-beta7. I build an observation from a double array \> new Observation(Nd4j.create(states)) where states is double of length 19. There is no problem while …

---

## [Out of Memory Error](https://community.konduit.ai/t/out-of-memory-error/466)

<div class="topic-metadata">

**Author:** [@haubna](https://community.konduit.ai/u/haubna)\
**Replies:** 8\
**Last updated:** [April 28, 2020, 11:14am UTC](https://community.konduit.ai/t/out-of-memory-error/466 "2020-04-28T11:14:51Z")

</div>

I am using RL4J QLearning and i constantly get an out of memory error. This is what my QLearning Configuration looks like, the memory leak doesn’t come from my application, I’ve tried giving it a ton of memory, tried Wor…

---

## [Custom MDP with RL4J](https://community.konduit.ai/t/custom-mdp-with-rl4j/64)

<div class="topic-metadata">

**Author:** [@diane](https://community.konduit.ai/u/diane)\
**Replies:** 1\
**Last updated:** [January 20, 2020, 2:16am UTC](https://community.konduit.ai/t/custom-mdp-with-rl4j/64 "2020-01-20T02:16:22Z")

</div>

How do I implement a custom MDP with RL4J?

---

## [Police-based algorithm](https://community.konduit.ai/t/police-based-algorithm/66)

<div class="topic-metadata">

**Author:** [@benjamin](https://community.konduit.ai/u/benjamin)\
**Replies:** 1\
**Last updated:** [January 20, 2020, 2:13am UTC](https://community.konduit.ai/t/police-based-algorithm/66 "2020-01-20T02:13:39Z")

</div>

About RL4j. To my knowledge, RL4j now support value-based rl algorithms like DQN(Double-DQN), any plan to develop policy-based algorithm ?
