Tokenize Thai sentence with ICUTokenizer JAVA

徘徊边缘 提交于 2019-12-11 07:18:02

问题


I am trying the below code to get all the tokens fro the thai sentence. It throws exception. Can anyone point me to tokenize thai in JAVA?

    import org.apache.lucene.analysis.Analyzer.TokenStreamComponents;
import org.apache.lucene.analysis.TokenFilter;
import org.apache.lucene.analysis.TokenStream;
import org.apache.lucene.analysis.icu.ICUNormalizer2Filter;
import org.apache.lucene.analysis.icu.segmentation.ICUTokenizer;
import org.apache.lucene.analysis.tokenattributes.CharTermAttribute;

public class Tokenizer{

    public static void main(String[] args) throws IOException {
        ICUTokenizer tokenizer = new ICUTokenizer(new StringReader("การที่ได้ต้องแสดงว่างานดี"));
        TokenFilter filter = new ICUNormalizer2Filter(tokenizer);
        TokenStreamComponents tt = new TokenStreamComponents(tokenizer, filter);
        TokenStream ts = tt.getTokenStream();
        CharTermAttribute cattr  = ts.addAttribute(CharTermAttribute.class);
        ts.reset();
        while(ts.incrementToken()){
            System.out.println(cattr.toString()+"-----");
        }
    }
}

Exception is as below

Exception in thread "main" java.lang.ExceptionInInitializerError
    at org.apache.lucene.analysis.icu.segmentation.ICUTokenizer.<init>(ICUTokenizer.java:72)
    at com.tokenizer.tt.main(tt.java:22)
Caused by: java.lang.RuntimeException: java.io.IOException: ICU data file error: Not an ICU data file
    at org.apache.lucene.analysis.icu.segmentation.DefaultICUTokenizerConfig.readBreakIterator(DefaultICUTokenizerConfig.java:128)
    at org.apache.lucene.analysis.icu.segmentation.DefaultICUTokenizerConfig.<clinit>(DefaultICUTokenizerConfig.java:66)
    ... 2 more
Caused by: java.io.IOException: ICU data file error: Not an ICU data file
    at com.ibm.icu.impl.ICUBinary.readHeader(ICUBinary.java:577)
    at com.ibm.icu.text.RBBIDataWrapper.get(RBBIDataWrapper.java:173)
    at com.ibm.icu.text.RuleBasedBreakIterator.getInstanceFromCompiledRules(RuleBasedBreakIterator.java:71)
    at org.apache.lucene.analysis.icu.segmentation.DefaultICUTokenizerConfig.readBreakIterator(DefaultICUTokenizerConfig.java:123)
    ... 3 more

回答1:


Finally figured out how to use ICU4J in a java program

import java.io.IOException;
import java.io.Reader;
import java.io.StringReader;
import org.apache.lucene.analysis.icu.segmentation.ICUTokenizer;
import org.apache.lucene.analysis.tokenattributes.CharTermAttribute;

public class icuEstes {

public static void main(String[] args) throws IOException {
    Reader reader = new StringReader("การที่ได้ต้องแสดงว่างานดี  This is a test ກວ່າດອກ");
    ICUTokenizer icut = new ICUTokenizer();
    icut.setReader(reader);
    icut.addAttribute(CharTermAttribute.class);
    icut.reset();
    while (icut.incrementToken()) {
        System.out.println(icut.toString());
        System.out.println(icut.getAttribute(CharTermAttribute.class));
    }
    icut.close();
}}


来源:https://stackoverflow.com/questions/43377330/tokenize-thai-sentence-with-icutokenizer-java

易学教程内所有资源均来自网络或用户发布的内容,如有违反法律规定的内容欢迎反馈
该文章没有解决你所遇到的问题?点击提问,说说你的问题,让更多的人一起探讨吧!