OCR For Indexing Scanned Articles– Google!

Oct 31, 2008 | 1,176 views | by Navneet Kaushal
VN:F [1.9.20_1166]
Rating: 0.0/5 (0 votes cast)

Earlier the user had to create a PDF(text based and not image based) in order to make a PDF file indexed by Google. Otherwise it was not possible for Googlebot to recognize the content.

According to an official announcement done by Google, the case is no longer the same.

Google says: “This Optical Character Recognition (OCR) technology lets us convert a picture (of a thousand words) into a thousand words — words that can be searched and indexed, so that these valuable documents are more easily found.

While we've indexed documents saved as PDFs for some time now, scanned documents are a lot more difficult for a computer to read. Scanning is the reverse of printing. Printing turns digital words into text on paper, while scanning makes a digital picture of the physical paper (and text) so you can store and view it on a computer. The scanned picture of the text is not quite the same as the original digital words, however — it is a picture of the printed words. Often you can see telltale signs: the ring of a coffee cup, ink smudges, or even fold creases in the pages.”

“To see our new system at work, click on these search queries. Note the document excerpt in the search results, along with the full text presented after the 'View as HTML' link:
[repairing aluminum wiring]
[spin lock performance]
[Mumps and Severe Neutropenia]
[Steady success in a volatile world]” -Google

After clicking on the first example, I found out that the first result is a Consumer Product Safety Commission PDF. It was very easily scanned as an image.

4.thumbnail OCR For Indexing Scanned Articles– Google!

Navneet Kaushal

Navneet Kaushal is the founder and CEO of PageTraffic, an SEO Agency in India with offices in Chicago, Mumbai and London. A leading search strategist, Navneet helps clients maintain an edge in search engines and the online media. Navneet's expertise has established PageTraffic as one of the most awarded and successful search marketing agencies.
4.thumbnail OCR For Indexing Scanned Articles– Google!
4.thumbnail OCR For Indexing Scanned Articles– Google!