因為這次案子的功能需求,在網路上找到兩種方式來分詞。一種是 SCWS 採用的詞庫方式,另一種是 WP 的 plugin Bigram Full-Text Search 所採用的 UTF-8 分詞方式。
詞庫方式的分詞系統大都採用知名部落客 蔡志浩 的論文:MMSEG: A Word Identification System for Mandarin Chinese Text Based on Two Variants of the Maximum Matching Algorithm 及 CScanner - A Chinese Lexical Scanner 中的理論來開發詞庫系統。
SCWS 簡易中文分詞系統 是由 hightman 開發,版權方面並沒有很清楚的宣告,僅僅宣告「本人保留一切相關權利」。但是套件中包含所有原始碼及正體/簡體中文詞庫,論壇中也有許多網友一起討論,就像一般 Open Source 套件開發社群的運作模式。另外也有其他網友把 scws 移植到 windows 環境。
Showing posts with label SCWS. Show all posts
Showing posts with label SCWS. Show all posts
November 17, 2008
Subscribe to:
Posts (Atom)
page top
Tags
3G
(1)
APT
(3)
baseball
(1)
blog
(1)
Blogger
(1)
cat
(9)
China
(2)
Cite
(1)
finance
(5)
Flickr
(1)
FreeTDS
(1)
GD
(1)
Google
(4)
Google Mail
(1)
Google Maps
(1)
Google News
(2)
hack
(1)
Hemidemi
(1)
history
(5)
JAVA
(4)
life
(1)
Lifetype
(3)
Linux
(1)
media
(5)
MOO
(1)
MS SQL
(1)
MSN
(1)
MySQL
(1)
Oracle
(1)
Peopo
(6)
photography
(2)
PHP
(4)
politics
(5)
Redhat
(3)
Resource Hacker
(1)
Roller
(1)
RSS
(2)
SCWS
(1)
SMTP
(1)
Sphinx
(1)
SVN
(1)
Tag
(1)
Taiwan
(7)
Tomcat
(4)
tour
(1)
Trademark
(1)
Ubuntu
(1)
video
(5)








