get the exact position of text from image in tesseract

前端 未结 2 1412
没有蜡笔的小新
没有蜡笔的小新 2021-02-01 09:36

Using GetHOCRText(0) method in tesseract I\'m able to retrieve the text in html and on presenting the html in webview i\'m able get the text but the postion of text in image is

相关标签:
2条回答
  • 2021-02-01 10:18

    GetBoxText() method will return exact position of each characters in an array.

    char *boxtext = _tesseract->GetBoxText(0);
    NSString* aBoxText = [NSString stringWithUTF8String:boxtext];
    
    0 讨论(0)
  • 2021-02-01 10:25

    If you have the hocr output, you should have a tag for each word. These tags should have class="ocrx_word" and name="bbox x1 y1 x2 y2" where the x and y are the top left and bottom right corner of the bounding box around the word. I don't think it's possible to automatically use this information to format a text document - would require translating pixel differences to number of tabs/spaces. But, you should be able to render text in the given location.

    0 讨论(0)
提交回复
热议问题