# Load Data Set Asynchronously From Data Base

**URL:** <https://community.konduit.ai/t/load-data-set-asynchronously-from-data-base/1612>\
**Category:** DL4J\
**Created:** [September 20, 2021, 1:48pm UTC](https://community.konduit.ai/t/load-data-set-asynchronously-from-data-base/1612 "2021-09-20T13:48:16Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![participant](https://avatars.discourse-cdn.com/v4/letter/p/94ad74/32.png) [@participant](https://community.konduit.ai/u/participant)\
**Post date:** [September 20, 2021, 1:48pm UTC](https://community.konduit.ai/t/load-data-set-asynchronously-from-data-base/1612/1 "2021-09-20T13:48:16Z")

</div>

Hi,  
I am trying to figure out a way to find a solution to what I think would be a common problem:

My training data is stored in a data base, and I want to avoid loading the full data set into memory. Also, I don’t need to preprocess the data anymore. (So I don’t really need datavec)

My approach would therefore be to implement a AsyncDataSetIterator, however, that is apparently not the way to go, as Adam Gibson pointed out here:

> <https://stackoverflow.com/questions/48845162/how-can-i-use-a-custom-data-model-with-deeplearning4j/49101775#49101775>

What would be the way to go here, or is a custom implementation of AsyncDataSetIterator indeed best to use here?

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [September 20, 2021, 2:58pm UTC](https://community.konduit.ai/t/load-data-set-asynchronously-from-data-base/1612/2 "2021-09-20T14:58:37Z")

</div>

In principle all you need to do is find or implement an appropriate RecordReader. DataSetIterators build from record readers by default load the data asynchronously.

As someone who has implemented several custom data set iterators, I can tell you that there are very few reasons to actually do it. And when you’re doing it, there are plenty of ways that you can mess it up.

As you’ve got your data in a database, I guess you will be able to connect to it via JDBC, so in principle you don’t even need a custom record reader, because a JDBC Record Reader already exists:

> <https://github.com/eclipse/deeplearning4j/blob/master/datavec/datavec-jdbc/src/main/java/org/datavec/jdbc/records/reader/impl/jdbc/JDBCRecordReader.java#L76-L86>

It isn’t exactly the best documented thing, but there are a few examples of using it in the tests:  
[https://github.com/eclipse/deeplearning4j/blob/master/datavec/datavec-jdbc/src/test/java/org/datavec/api/records/reader/impl/JDBCRecordReaderTest.java#L307-L311](https://github.com/eclipse/deeplearning4j/blob/master/datavec/datavec-jdbc/src/test/java/org/datavec/api/records/reader/impl/JDBCRecordReaderTest.java#L307-L311)

To create a dataset iterator from it, you use the same [RecordReaderDataSetIterator](https://javadoc.io/doc/org.deeplearning4j/deeplearning4j-datavec-iterators/latest/org/deeplearning4j/datasets/datavec/RecordReaderDataSetIterator.Builder.html) as with any other record reader and you’ll get a well working iterator without all the headache.
