# Finding a word in an image

**URL:** <https://community.glideapps.com/t/finding-a-word-in-an-image/55634>\
**Category:** Ask for Help\
**Created:** [December 10, 2022, 5:23pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634 "2022-12-10T17:23:13Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![AyS\_0908](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/ays_0908/32/27349_2.png) [@AyS\_0908](https://community.glideapps.com/u/AyS_0908)\
**Post date:** [December 10, 2022, 5:23pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/1 "2022-12-10T17:23:13Z")

</div>

Hi,

I would like to build a very basic app (from Glide point of view) but with one “little detail” that I have absolutely no idea how to start it.

My use case is that I have a shelf with 300 books. And I never find the one I am looking for.  
So, I would like to:

- take a picture of the shelf
- enter the title of the book to be found
- then have the image modified with a frame on the title of the book.

 ![image](https://us1.discourse-cdn.com/flex002/uploads/glideapps/original/3X/2/5/25d636e8d8cc8ad625903967a235b7fa2f57d107.png)

If you had some advices, it would be great (knowing that my knowledge in code and AI is below 0).

Thanks in advance

---

<div class="post-metadata">

**Author:** ![Robert\_Petitto](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/robert_petitto/32/25193_2.png) [@Robert\_Petitto](https://community.glideapps.com/u/Robert_Petitto)\
**Post date:** [December 10, 2022, 8:08pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/2 "2022-12-10T20:08:03Z")

</div>

Frame around the book will be tricky, but an arrow pointing to the book should be much easier.

Something like this:

1. Upload photo to Glide
2. Find book button → Send image and keyword to Zapier/Make → Analyze with Google Cloud Vision → Find X/Y of located text → send coordinates back to Glide → cloudinary toadd arrow overlay pointing at X/Y

---

<div class="post-metadata">

**Author:** ![AyS\_0908](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/ays_0908/32/27349_2.png) [@AyS\_0908](https://community.glideapps.com/u/AyS_0908)\
**Post date:** [December 10, 2022, 10:14pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/3 "2022-12-10T22:14:18Z")

</div>

An arrow will completely ok, I just want to identify where it is.  
Thanks @Robert_Petitto … let’s discover Google Cloud Vision

(“nocode” they used to say 🙂 )

---

<div class="post-metadata">

**Author:** ![ThinhDinh](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/thinhdinh/32/49_2.png) [@ThinhDinh](https://community.glideapps.com/u/ThinhDinh)\
**Post date:** [December 11, 2022, 12:21am UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/4 "2022-12-11T00:21:30Z")

</div>

I imagine if the front of the book is visible to the computer vision then it would be easier.

 ![image](https://us1.discourse-cdn.com/flex002/uploads/glideapps/original/3X/b/e/be37dc01f9230658b5f6532b898af824ab5cabd3.jpeg)

But if you have something like this, which I expect to be the standard, that might be a little harder to read the text out, since they are presented vertically.

---

<div class="post-metadata">

**Author:** ![AyS\_0908](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/ays_0908/32/27349_2.png) [@AyS\_0908](https://community.glideapps.com/u/AyS_0908)\
**Post date:** [December 11, 2022, 9:44am UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/5 "2022-12-11T09:44:18Z")

</div>

Yes @ThinhDinh this is exactly what I want to build…

I was thinking about having an option copying three time the original image then turning one 90°, one -90° and one 180°.  
And finally send them three to computer vision.

But this was before going to “google cloud vision”… nightmare, don’t even know where to start…

nb: just for fun, I asked to ChatGPT, it recommends tensorflow (?) and gives the code!

---

<div class="post-metadata">

**Author:** ![Alvaro\_Souza](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/alvaro_souza/32/51490_2.png) [@Alvaro\_Souza](https://community.glideapps.com/u/Alvaro_Souza)\
**Post date:** [December 11, 2022, 3:58pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/6 "2022-12-11T15:58:45Z")

</div>

Hi!

i dont know if im late for this, but a solution could be using a no-code web app that use data science templates for idetifying what you are looking for.

for example, using this web app: [“obviously.ai”](https://obviously.ai/) can already do what you want to, but like every data science project you need data to make it better (in your case a lot of images and photos)

then connect the results to glide with any web automation (like make, zapier or power automate) and voila!

i hope it help you! 🙂

---

<div class="post-metadata">

**Author:** ![AyS\_0908](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/ays_0908/32/27349_2.png) [@AyS\_0908](https://community.glideapps.com/u/AyS_0908)\
**Post date:** [December 11, 2022, 5:04pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/7 "2022-12-11T17:04:41Z")

</div>

Thank you @Alvaro_Souza for your answer.

I have 0 knowledge in AI, and have a basic question: given my app should “just” read some typo letters, cf. book title (which would be previously entered in a form), why do I need to train a model? Is letters recognition not something quite basic?

 ![image](https://us1.discourse-cdn.com/flex002/uploads/glideapps/original/3X/b/6/b6fc29c3fbc556a8a0f876168b69bdbec95a32f0.png)

Thanks in advance

---

<div class="post-metadata">

**Author:** ![Alvaro\_Souza](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/alvaro_souza/32/51490_2.png) [@Alvaro\_Souza](https://community.glideapps.com/u/Alvaro_Souza)\
**Post date:** [December 12, 2022, 12:39am UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/8 "2022-12-12T00:39:31Z")

</div>

as i undesrtand what you need, the model have to first recognize that is a book and then read whats on them. right?

---

<div class="post-metadata">

**Author:** ![AyS\_0908](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/ays_0908/32/27349_2.png) [@AyS\_0908](https://community.glideapps.com/u/AyS_0908)\
**Post date:** [December 12, 2022, 9:35am UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/9 "2022-12-12T09:35:14Z")

</div>

Well, I’m not sure that it’s mandatory to recognize that it is a book.  
The goal is to find a “title”.  
Therefore, on below picture for example, it " “just” " has to search for text and show where it matches the title looked for

 ![image](https://us1.discourse-cdn.com/flex002/uploads/glideapps/original/3X/e/f/ef12bf4233f590f73860c53acc6d8433a67453cf.jpeg)

---

<div class="post-metadata">

**Author:** ![Jeff\_Hager](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/jeff_hager/32/43_2.png) [@Jeff\_Hager](https://community.glideapps.com/u/Jeff_Hager)\
**Post date:** [December 12, 2022, 12:44pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/10 "2022-12-12T12:44:22Z")

</div>

Google Lens is pretty good at reading text in any orientation. I use it a lot for translating other languages in Images, or live in real time with my camera. Also for copying text out of images. Pretty good for identifying plants too. It even solves typed or handwritten math problems. Not sure though if they have an API or any way to use it as a service. Pretty powerful feature on my phone though.

I found a random picture of books online and with Google Lens it was about to find most of the book titles. It’s subtle, but all text is highlighted in this image and I can click on each book to do a Google search or copy that text.

 ![Screenshot_20221212-063042](https://us1.discourse-cdn.com/flex002/uploads/glideapps/original/3X/c/f/cf256b3a01f01bb653d36eef8ae5291a62dc55e3.jpeg)

I can even translate everything to French.

 ![Screenshot_20221212-064306](https://us1.discourse-cdn.com/flex002/uploads/glideapps/original/3X/4/2/42da4b38ae06f8840e2ba17f10d43e5aac0cf09d.jpeg)

If you can find something that uses Google Lens in some way, then that might work for you.

---

<div class="post-metadata">

**Author:** ![Robert\_Petitto](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/robert_petitto/32/25193_2.png) [@Robert\_Petitto](https://community.glideapps.com/u/Robert_Petitto)\
**Post date:** [December 12, 2022, 1:02pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/11 "2022-12-12T13:02:30Z")

</div>

Ya…Google Lens = Google Cloud Vision. [Make](https://www.make.com/en/register?pc=robert) has a scenario for that.

---

<div class="post-metadata">

**Author:** ![AyS\_0908](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/ays_0908/32/27349_2.png) [@AyS\_0908](https://community.glideapps.com/u/AyS_0908)\
**Post date:** [December 12, 2022, 1:54pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/12 "2022-12-12T13:54:36Z")

</div>

Thank you both,  
I’ve tried to plays with Google Cloud Vision, it’s complicated for a beginner… I’m still trying to generate an API Key 🙂  
[edit: I’ve asked to ChatGPT and it explained me the how to!!]

But what @Jeff_Hager achieved is exactly what I will try to do (+ a form to enter the “title” searched)

[re: edit] I achieved to initiate this “star wars nocode” suite!  
I “just” have to have the “blue” part to do something…🙃😵‍💫 because up to now it just copies-pastes the initial picture

 ![image](https://us1.discourse-cdn.com/flex002/uploads/glideapps/original/3X/f/7/f7ca47a63ca8c72bf3719305db107d676a453043.png)

---

<div class="post-metadata">

**Author:** ![Robert\_Petitto](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/robert_petitto/32/25193_2.png) [@Robert\_Petitto](https://community.glideapps.com/u/Robert_Petitto)\
**Post date:** [December 12, 2022, 11:34pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/13 "2022-12-12T23:34:27Z")

</div>

The response to the Google cloud vision module should return XY coordinates of the text it’s instructed to find. If you’re able to achieve this, let me know and I can give you next steps.

---

<div class="post-metadata">

**Author:** ![AyS\_0908](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/ays_0908/32/27349_2.png) [@AyS\_0908](https://community.glideapps.com/u/AyS_0908)\
**Post date:** [December 21, 2022, 4:36pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/14 "2022-12-21T16:36:36Z")

</div>

Hi, as @Robert_Petitto said, Google Vision + Make.

**BUT** what Robert did not say 🙂 : there is a 🐍 → you need to do some Python !, or java, go and all their barbarian friends!  
(cf. "Label detection part)

I could copy the code, play loto by changing some lines, but I don’t even know where pasting it 🙂  
I think it is a step too high a poor lonesome NO-coder !

> **[Quickstart: Detect labels in an image by using client libraries  | ...](https://cloud.google.com/vision/docs/detect-labels-image-client-libraries)**
>
> Learn how to detect labels in an image by using Vision API in your favorite programming language.

---

<div class="post-metadata">

**Author:** ![slscustom.ru](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/slscustom.ru/32/77336_2.png) [@slscustom.ru](https://community.glideapps.com/u/slscustom.ru)\
**Post date:** [December 24, 2022, 8:39pm UTC](https://community.glideapps.com/t/finding-a-word-in-an-image/55634/15 "2022-12-24T20:39:52Z")

</div>

Hi! Im not clear about topic, but i remember cloudinary have addons to recognize text on images. No code solution.
