# Referencing 100k+ rows lists

**URL:** <https://community.glideapps.com/t/referencing-100k-rows-lists/78935>\
**Category:** Ask for Help\
**Tags:** api, actions\
**Created:** [January 2, 2025, 8:30pm UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935 "2025-01-02T20:30:48Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mikaeldorval](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/mikaeldorval/32/67405_2.png) [@Mikaeldorval](https://community.glideapps.com/u/Mikaeldorval)\
**Post date:** [January 2, 2025, 8:30pm UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/1 "2025-01-02T20:30:48Z")

</div>

I’m building a quoting system that adds up the total amount of government grant depending on the quoted items.

Grants amounts depend on the model numbers used in different combinations.

I need the system to lookup a few pdf docs available on government websites and compare the quotes data to match the exact grants amounts.

Those pdf docs have up to 100k+ rows of data.

I tried:

- Converting docs to text and extracting the info with AI  
Result: String too long

- Looking up doc link with AI  
Result: Request too big

I’m looking for creative solutions. I tried AI with shorter sample of the same docs and it works great but I can’t import 100k+ rows in my Glide app just for lookup purposes.

Any ideas?

---

<div class="post-metadata">

**Author:** ![MaximeBaker](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/maximebaker/32/80877_2.png) [@MaximeBaker](https://community.glideapps.com/u/MaximeBaker)\
**Post date:** [January 2, 2025, 10:54pm UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/2 "2025-01-02T22:54:31Z")

</div>

> [@Mikaeldorval](#):
>
> I need the system to lookup a few pdf docs available on government websites and compare the quotes data to match the exact grants amounts.

How do you get these files?

> [@Mikaeldorval](#):
>
> Those pdf docs have up to 100k+ rows of data.

You meant pages?

> [@Mikaeldorval](#):
>
> - Converting docs to text and extracting the info with AI  
> Result: String too long
> - Looking up doc link with AI  
> Result: Request too big

You can try to use an external workflow to split the PDFs into multiple chunks, then use these chunks in the AI action/column.

I think it might be possible with Make, but I am not familiar with Make. I used Make maybe twice in my life 😂

---

<div class="post-metadata">

**Author:** ![Mikaeldorval](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/mikaeldorval/32/67405_2.png) [@Mikaeldorval](https://community.glideapps.com/u/Mikaeldorval)\
**Post date:** [January 2, 2025, 11:52pm UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/3 "2025-01-02T23:52:34Z")

</div>

I get the links for each list here: [Archive des listes des thermopompes efficaces admissibles | Hydro-Québec](https://www.hydroquebec.com/residentiel/mieux-consommer/aides-financieres/logisvert/archives-liste-thermopompes.html)

The first one is 3092 pages and ~40 rows per page.

---

<div class="post-metadata">

**Author:** ![MaximeBaker](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/maximebaker/32/80877_2.png) [@MaximeBaker](https://community.glideapps.com/u/MaximeBaker)\
**Post date:** [January 3, 2025, 1:29am UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/5 "2025-01-03T01:29:18Z")

</div>

Is there any online table version of these datas?

---

<div class="post-metadata">

**Author:** ![Mikaeldorval](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/mikaeldorval/32/67405_2.png) [@Mikaeldorval](https://community.glideapps.com/u/Mikaeldorval)\
**Post date:** [January 3, 2025, 2:52am UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/6 "2025-01-03T02:52:17Z")

</div>

Not for now…

---

<div class="post-metadata">

**Author:** ![MaximeBaker](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/maximebaker/32/80877_2.png) [@MaximeBaker](https://community.glideapps.com/u/MaximeBaker)\
**Post date:** [January 3, 2025, 3:22am UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/7 "2025-01-03T03:22:22Z")

</div>

Did you try with google vision ?

---

<div class="post-metadata">

**Author:** ![Mikaeldorval](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/mikaeldorval/32/67405_2.png) [@Mikaeldorval](https://community.glideapps.com/u/Mikaeldorval)\
**Post date:** [January 3, 2025, 3:42am UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/8 "2025-01-03T03:42:14Z")

</div>

I just tried it, still returns an error

---

<div class="post-metadata">

**Author:** ![MaximeBaker](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/maximebaker/32/80877_2.png) [@MaximeBaker](https://community.glideapps.com/u/MaximeBaker)\
**Post date:** [January 3, 2025, 1:09pm UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/9 "2025-01-03T13:09:34Z")

</div>

I think that the best solution would be to split the PDFs…

---

<div class="post-metadata">

**Author:** ![ThinhDinh](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/thinhdinh/32/49_2.png) [@ThinhDinh](https://community.glideapps.com/u/ThinhDinh)\
**Post date:** [January 4, 2025, 1:24am UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/10 "2025-01-04T01:24:38Z")

</div>

So splitting the PDF and use regex to find the relevant info, right?

---

<div class="post-metadata">

**Author:** ![MaximeBaker](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/maximebaker/32/80877_2.png) [@MaximeBaker](https://community.glideapps.com/u/MaximeBaker)\
**Post date:** [January 4, 2025, 2:53am UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/11 "2025-01-04T02:53:46Z")

</div>

Or just splitting the PDFs into chunks and use them all.

---

<div class="post-metadata">

**Author:** ![Mikaeldorval](https://sea2.discourse-cdn.com/flex002/user_avatar/community.glideapps.com/mikaeldorval/32/67405_2.png) [@Mikaeldorval](https://community.glideapps.com/u/Mikaeldorval)\
**Post date:** [January 4, 2025, 4:23pm UTC](https://community.glideapps.com/t/referencing-100k-rows-lists/78935/12 "2025-01-04T16:23:01Z")

</div>

I could split the pdf using make but never used regex. I’ll look into that. Thanks for the info!
