Algorithm to detect similar documents in python script [closed]

前端未结

关注

 10  1519

时光说笑

相关标签:

10条回答

天涯浪人

2020-12-24 04:20

If these are pure text documents, or you have a method to extract the text from the documents, you can use a technique called shingling.

You first compute a unique hash for each document. If these are the same, you are done.

If not, you break each document down into smaller chunks. These are your 'shingles.'

Once you have the shingles, you can then compute identity hashes for each shingle and compare the hashes of the shingles to determine if the documents are actually the same.

The other technique you can use is to generate n-grams of the entire documents and compute the number of similar n-grams in each document and produce a weighted score for each document. Basically an n-gram is splitting a word into smaller chunks. 'apple' would become ' a', ' ap', 'app', 'ppl', 'ple', 'le '. (This is technically a 3-gram) This approach can become quite computationally expensive over a large number of documents or over two very large documents. Of course, common n-grams 'the', ' th, 'th ', etc need to be weighted to score them lower.

I've posted about this on my blog and there are some links in the post to a few other articles on the subject Shingling - it's not just for roofers.

Best of luck!

0 讨论(0)
发布评论:

提交评论
- 加载中...
情书的邮戳

2020-12-24 04:22

I think Jeremy has hit the nail on the head - if you just want to detect if files are different, a hash algorithm like MD5 or SHA1 is a good way to go.

Linus Torvalds' Git source control software uses SHA1 hashing in just this way - to check when files have been modified.

0 讨论(0)
发布评论:

提交评论
- 加载中...
忘了有多久

2020-12-24 04:22

You might want to look into the DustBuster algorithm as outlined in this paper.

From the paper, they're able to detect duplicate pages without even examining the page contents. Of course examining the contents increases the efficacy, but using raw server logs is adequate for the method to detect duplicate pages.

Similar to the recommendation of using MD5 or SHA1 hashes, the DustBuster method largely relies on comparing file size as it primary signal. As simple as it sounds, it's rather effective for an initial first pass.

0 讨论(0)
发布评论:

提交评论
- 加载中...
滥情空心

2020-12-24 04:24
You can use or at last study difflib from Python's stdlib to write your code.

It is very flexible, and has algorithms to find differences between lists of strings, and to point these differences. Then you can use the get_close_matches() to find similar words:
```
>>> get_close_matches('appel', ['ape', 'apple', 'peach', 'puppy'])
['apple', 'ape']
```
It is not the solution but maybe it is a start.
0 讨论(0)
发布评论:

提交评论
- 加载中...

我在风中等你

2020-12-24 04:26

Similarity can be found easily without classification. Try this O(n2) but works fine.

def jaccard_similarity(doc1, doc2):
    a = sets(doc1.split())
    b = sets(doc2.split())
    similarity = float(len(a.intersection(b))*1.0/len(a.union(b))) #similarity belongs to [0,1] 1 means its exact replica.
    return similarity

0 讨论(0)

梦谈多话

2020-12-24 04:26

There is a rather good talk on neural networks on Google Techtalks that talks about using layered Boltzmann machines to generate feature vectors for documents that can then be used to measure document distance. The main issue is the requirement to have a large sample document set to train the network to discover relevant features.

0 讨论(0)
发布评论:

提交评论
- 加载中...

1 2 下一页

热议问题