undefined

points

[-]

Did you take any steps to decrease the dimension size of images, if this increases the performance? I have not tried this as I have not peformed an OCR task like this with an LLM. I would be interested to know at what size the vlm cannot make out the details in text reliably.

by embedding-shape8 hours ago|

parent|

[-]

The performance is OK, takes a couple of seconds at most on my GPU, just the amount of documents to get through that takes time, even with parallelism. The dimension seems fine as it is, as far as I can tell.

by helterskelter8 hours ago|

prev|

[-]

[flagged]

by embedding-shape8 hours ago|

parent|

[-]

Haven't seen anything particular about that, but lots of the documents with names that were half-redacted contain OCRd text that is completely garbled, but olmocr-2-7b seems to handle it just fine. Unsure if they just had sucky processes or if there is something else going on.

by helterskelter8 hours ago|

parent|

[-]

Might be a good fit for uploading a git repo and crowdsourcing

by embedding-shape7 hours ago|

parent|

[-]

Was my first impulse too but not sure I trust that unless I could gather a bunch of people I trust, which would mean I'd no longer be anonymous. Kind of a catch22.

by direwolf202 hours ago|

parent|

prev|

[-]

GitHub would ban you