# OCR in deeplearning4j

**URL:** <https://community.konduit.ai/t/ocr-in-deeplearning4j/197>\
**Category:** DL4J\
**Created:** [March 1, 2020, 4:55pm UTC](https://community.konduit.ai/t/ocr-in-deeplearning4j/197 "2020-03-01T16:55:26Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![powroseba](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/powroseba/32/73_2.png) [@powroseba](https://community.konduit.ai/u/powroseba)\
**Post date:** [March 1, 2020, 4:55pm UTC](https://community.konduit.ai/t/ocr-in-deeplearning4j/197/1 "2020-03-01T16:55:26Z")

</div>

Hello,

I am trying to find working example of _optical character recognition_ in deeplearning4j without any result. Have you ever seen that example or could you give me some advice how to prepare data to learn network to learn it predict text in image ? Thanks in advance for help

---

<div class="post-metadata">

**Author:** ![bewithme](https://avatars.discourse-cdn.com/v4/letter/b/41988e/32.png) [@bewithme](https://community.konduit.ai/u/bewithme)\
**Post date:** [March 2, 2020, 5:53am UTC](https://community.konduit.ai/t/ocr-in-deeplearning4j/197/2 "2020-03-02T05:53:44Z")

</div>

please reference this  
demo org.deeplearning4j.examples.convolution.captcharecognition.MultiDigitNumberRecognition

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [March 2, 2020, 1:29pm UTC](https://community.konduit.ai/t/ocr-in-deeplearning4j/197/3 "2020-03-02T13:29:08Z")

</div>

While the [MultiDigitNumberRecognition](https://github.com/eclipse/deeplearning4j-examples/blob/master/dl4j-examples/src/main/java/org/deeplearning4j/examples/convolution/captcharecognition/MultiDigitNumberRecognition.java) is a good starting point, you should also remember that OCR in general isn’t an easy problem.

So you might also want to read about the tesseract project (e.g. [docs/das\_tutorial2016 at main · tesseract-ocr/docs · GitHub](https://github.com/tesseract-ocr/docs/tree/master/das_tutorial2016)) and how they are training an lstm for OCR: [https://github.com/tesseract-ocr/tessdoc/blob/master/TrainingTesseract-4.00.md](https://github.com/tesseract-ocr/tessdoc/blob/master/TrainingTesseract-4.00.md)

---

<div class="post-metadata">

**Author:** ![saudet](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/saudet/32/69_2.png) [@saudet](https://community.konduit.ai/u/saudet)\
**Post date:** [March 2, 2020, 2:44pm UTC](https://community.konduit.ai/t/ocr-in-deeplearning4j/197/4 "2020-03-02T14:44:59Z")

</div>

With wrappers for Java available here:

> **[javacpp-presets/tesseract at master · bytedeco/javacpp-presets](https://github.com/bytedeco/javacpp-presets/tree/master/tesseract)**
>
> master/tesseract
