1スキャンした PDF documents をアップロードします。クリーンな 300 DPI グレースケールスキャンは認識機能にとって最も有用です。
2言語を指定すると、 言語の名前を付けることで、 推測をさせることよりも 多くの利点が得られます。
3認識パスを実行します。単語は元のページ画像の上に見えないレイヤーとして書き戻されます。
4検索可能な PDF ファイルをダウンロードします。同じように見えますが、検索とテキストの選択が可能です。
OCR PDF よくある質問
ソースフォーマットは認識に影響するのか?
+
It does. PDF はページを固定します: フォント、ベクトル、ラスター画像、テキスト座標は凍結され、読者全員が同じレイアウトを見るようになります. How the page is stored decides what resolution and colour information the recogniser has to work with.
PDF に特有な何かは?
+
Yes — PDFはページの画像ではなく オブジェクトグラフです テキストは選択可能で ベクトルは鋭く残っています 内部のラスターコンテンツに何が起こっても. It affects what the recogniser can see.
OCR PDF はどうやってスキャンをテキストに変換しますか?
+
Concretely, 各ページはレンダリングされ テッセラクトのテキスト認識を通して 100以上の言語で実行され 認識された単語は 原画像の上に 見えないテキストレイヤーとして書き戻されます ページは変わらないように見えますが 検索可能で選択可能です. The page still looks exactly as it did — the recognised text sits invisibly behind the image so search and selection work without changing the appearance.
WEBP.toは、旧来のデフォルトの両方の現代的な置き換えを中心に構築されている:損失、無損失、アルファ、アニメーションを行うフォーマット、そして、JPEGやPNGよりも信頼性の高い小さなサイズに到達する。 People arrive here having decided the old defaults are costing them bandwidth, and that decision immediately raises the rest: what size, what quality, what happens to the transparency. OCR PDF is on the same upload and the same account so those questions get answered in one place.
OCR PDF が実行されたら、次に何をする価値があるか?
+
このサイトのコンバータは、WebP、PNG、JPG、AVIFの画像を移動します。小さなファイルが古いブラウザを失う価値があるのかどうかを決めるのです。 Doing it in that order matters: pick the pixels first and the container afterwards, because the container is the cheap decision and the pixels are the expensive one.