Jieba

jieba 토큰화 도구는 중국어 텍스트를 구성 단어로 분해하여 처리합니다.

jieba 토큰화 도구는 출력에서 구두점을 별도의 토큰으로 보존합니다. 예를 들어 "你好!世界。" 은 ["你好", "!", "世界", "。"] 이 됩니다. 이러한 독립형 구두점 토큰을 제거하려면 removepunct 필터를 사용합니다.

구성

Milvus는 jieba 토큰화기에 대해 단순 구성과 사용자 정의 구성이라는 두 가지 구성 방식을 지원합니다.

단순 구성

단순 구성의 경우 토큰화기를 "jieba" 로 설정하기만 하면 됩니다. 예를 들어

# Simple configuration: only specifying the tokenizer name
analyzer_params = {
    "tokenizer": "jieba",  # Use the default settings: dict=["_default_"], mode="search", hmm=True
}
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("tokenizer", "jieba");
const analyzer_params = {
    "tokenizer": "jieba",
};
analyzerParams = map[string]any{"tokenizer": "jieba"}
# restful
analyzerParams='{
  "tokenizer": "jieba"
}'

이 간단한 구성은 다음 사용자 지정 구성과 동일합니다:

# Custom configuration equivalent to the simple configuration above
analyzer_params = {
    "type": "jieba",          # Tokenizer type, fixed as "jieba"
    "dict": ["_default_"],     # Use the default dictionary
    "mode": "search",          # Use search mode for improved recall (see mode details below)
    "hmm": True                # Enable HMM for probabilistic segmentation
}
Map<String, Object> analyzerParams = new HashMap<>();
analyzerParams.put("type", "jieba");
analyzerParams.put("dict", Collections.singletonList("_default_"));
analyzerParams.put("mode", "search");
analyzerParams.put("hmm", true);
// javascript
analyzerParams = map[string]any{"type": "jieba", "dict": []any{"_default_"}, "mode": "search", "hmm": true}
# restful

매개변수에 대한 자세한 내용은 사용자 지정 구성을 참조하세요.

사용자 지정 구성

더 많은 제어를 위해 사용자 지정 사전을 지정하고, 세분화 모드를 선택하고, 숨겨진 마르코프 모델(HMM)을 활성화 또는 비활성화할 수 있는 사용자 지정 구성을 제공할 수 있습니다. 예를 들어

# Custom configuration with user-defined settings
analyzer_params = {
    "tokenizer": {
        "type": "jieba",           # Fixed tokenizer type
        "dict": ["customDictionary"],  # Custom dictionary list; replace with your own terms
        "mode": "exact",           # Use exact mode (non-overlapping tokens)
        "hmm": False               # Disable HMM; unmatched text will be split into individual characters
    }
}
Map<String, Object> analyzerParams = new HashMap<>();                                                                          
analyzerParams.put("tokenizer", new HashMap<String, Object>() {{
  put("type", "jieba");                                                                                                      
  put("dict", Arrays.asList("customDictionary"));             
  put("mode", "exact");
  put("hmm", false);
}});

// javascript
analyzerParams := map[string]interface{}{
  "tokenizer": map[string]interface{}{
      "type": "jieba",
      "dict": []string{"customDictionary"},
      "mode": "exact",
      "hmm":  false,
  },
}
# restful

파라미터

설명

기본값

type

토큰화기의 유형입니다. "jieba" 로 고정되어 있습니다.

"jieba"

dict

분석기가 어휘 소스로 로드할 사전 목록입니다. 기본 제공 옵션:

  • "_default_": 엔진에 내장된 중국어 간체 사전을 로드합니다. 자세한 내용은 dict.txt를 참조하세요.

  • "_extend_default_": "_default_" 의 모든 항목과 추가 중국어 번체 보충 자료를 로드합니다. 자세한 내용은 dict.txt.big을 참조하세요.

    기본 제공 사전과 사용자 지정 사전을 얼마든지 혼합할 수도 있습니다. 예: ["_default_", "结巴分词器"].

["_default_"]

mode

세분화 모드. 사용 가능한 값

  • "exact": 가장 정확한 방식으로 문장을 세분화하여 텍스트 분석에 이상적입니다.

  • "search": 정확한 모드를 기반으로 긴 단어를 더 세분화하여 기억력을 향상시켜 검색 엔진 토큰화에 적합합니다.

    자세한 내용은 Jieba 깃허브 프로젝트를 참조하세요.

"search"

hmm

사전에서 찾을 수 없는 단어의 확률적 세분화를 위해 숨겨진 마르코프 모델(HMM)을 활성화할지 여부를 나타내는 부울 플래그입니다.

true

dict 을 통해 인라이닝하는 대신 외부 파일에서 대용량 사용자 정의 어휘를 로드하려면 아래의 사전 파일을 사용한 사용자 정의 구성을 참조하세요.

analyzer_params 을 정의한 후 컬렉션 스키마를 정의할 때 VARCHAR 필드에 적용할 수 있습니다. 이렇게 하면 Milvus가 효율적인 토큰화 및 필터링을 위해 지정된 분석기를 사용하여 해당 필드의 텍스트를 처리할 수 있습니다. 자세한 내용은 사용 예시를 참조하세요.

사전 파일을 사용한 사용자 지정 구성Compatible with Milvus 3.0.x

도메인 용어집, 제품 용어집 또는 고유 명사 목록과 같은 대규모 사용자 정의 어휘의 경우, 단어를 파일에 저장하고 파일을 원격 파일 리소스로 등록한 다음 extra_dict_file 매개 변수를 통해 토큰화기에서 해당 파일을 참조하세요. 분석기는 이러한 단어를 기본 제공 사전 위에 있는 어휘집에 로드합니다.

파일은 한 줄에 하나의 용어가 포함된 일반 UTF-8 텍스트입니다. 예를 들어

结巴分词器
向量数据库

Milvus 클러스터가 사용하도록 구성된 개체 저장소에 파일을 업로드한 다음 등록합니다:

from pymilvus import MilvusClient

client = MilvusClient(uri="http://localhost:19530")

# Register the uploaded file under a name you'll reference from analyzer configs.
client.add_file_resource(
    name="zh_terms",
    path="file/zh_terms.txt",    # full S3 object key, including the rootPath prefix
)
// java
// nodejs
// go
# restful

extra_dict_file 을 통해 토큰화기에서 등록된 리소스를 참조합니다:

analyzer_params = {
    "tokenizer": {
        "type": "jieba",
        "dict": ["_default_"],             # keep the built-in dictionary
        "mode": "exact",
        "hmm": False,
        "extra_dict_file": {
            "type": "remote",
            "resource_name": "zh_terms",
            "file_name": "zh_terms.txt",
        },
    },
}

client.run_analyzer(["milvus结巴分词器中文测试"], analyzer_params)
# → [['milvus', '结巴', '分词器', '中文', '测试']]
// java
// nodejs
// go
# restful

extra_dict_file 파라미터는 다음과 같은 필드를 가진 객체를 허용합니다:

필드

설명

type

리소스 유형. add_file_resource 을 통해 등록된 파일의 경우 "remote" 을 사용합니다. 자체 호스팅 배포에 사용되는 "local" 변형에 대해서는 파일 리소스 관리를 참조하세요.

resource_name

파일이 add_file_resource 에 등록될 때 사용된 이름입니다.

file_name

등록된 리소스의 객체 저장소 경로 중 파일 이름 부분(예: 리소스가 path="file/zh_terms.txt" 에 등록된 경우 "zh_terms.txt" )입니다.

extra_dict_file 을 통해 추가된 단어는 기본 제공 사전과 병합되므로 jieba의 세분화 알고리즘은 기존 항목과 함께 해당 단어를 보게 됩니다. 특정 용어가 독립형 토큰으로 표시되는지 여부는 jieba의 확률 가중치 DAG 선택에 따라 달라집니다. 向量数据库 같은 긴 사용자 정의 용어는 기본 제공 사전에서 더 짧은 항목의 빈도가 더 높은 경우 向量 + 数据库 로 분할될 수 있습니다.

예제

분석기 구성을 컬렉션 스키마에 적용하기 전에 run_analyzer 메서드를 사용하여 그 동작을 확인하세요.

분석기 구성

analyzer_params = {
    "tokenizer": {
        "type": "jieba",
        "dict": ["结巴分词器"],
        "mode": "exact",
        "hmm": False
    }
}
Map<String, Object> analyzerParams = new HashMap<>();                                                                          
analyzerParams.put("tokenizer", new HashMap<String, Object>() {{
  put("type", "jieba");                                                                                                      
  put("dict", Arrays.asList("结巴分词器"));                   
  put("mode", "exact");
  put("hmm", false);
}});
// javascript
analyzerParams := map[string]interface{}{
  "tokenizer": map[string]interface{}{
      "type": "jieba",
      "dict": []string{"结巴分词器"},
      "mode": "exact",
      "hmm":  false,
  },
}
# restful

다음을 사용하여 확인 run_analyzer

from pymilvus import (
    MilvusClient,
)

client = MilvusClient(
    uri="http://localhost:19530",
    token="root:Milvus"
)

# Sample text to analyze
sample_text = "milvus结巴分词器中文测试"

# Run the standard analyzer with the defined configuration
result = client.run_analyzer(sample_text, analyzer_params)
print("Standard analyzer output:", result)
import io.milvus.v2.client.ConnectConfig;
import io.milvus.v2.client.MilvusClientV2;
import io.milvus.v2.service.vector.request.RunAnalyzerReq;
import io.milvus.v2.service.vector.response.RunAnalyzerResp;

ConnectConfig config = ConnectConfig.builder()
        .uri("http://localhost:19530")
        .token("root:Milvus")
        .build();
MilvusClientV2 client = new MilvusClientV2(config);

List<String> texts = new ArrayList<>();
texts.add("milvus结巴分词器中文测试");

RunAnalyzerResp resp = client.runAnalyzer(RunAnalyzerReq.builder()
        .texts(texts)
        .analyzerParams(analyzerParams)
        .build());
List<RunAnalyzerResp.AnalyzerResult> results = resp.getResults();
// javascript
import (
    "context"
    "encoding/json"
    "fmt"

    "github.com/milvus-io/milvus/client/v2/milvusclient"
)

client, err := milvusclient.New(ctx, &milvusclient.ClientConfig{
    Address: "localhost:19530",
    APIKey:  "root:Milvus",
})
if err != nil {
    fmt.Println(err.Error())
    // handle error
}

bs, _ := json.Marshal(analyzerParams)
texts := []string{"milvus结巴分词器中文测试"}
option := milvusclient.NewRunAnalyzerOption(texts).
    WithAnalyzerParams(string(bs))

result, err := client.RunAnalyzer(ctx, option)
if err != nil {
    fmt.Println(err.Error())
    // handle error
}
# restful

예상 출력

['milvus', '结巴分词器', '中', '文', '测', '试']

번역DeepL

관리형 Milvus를 무료로 사용해 보세요

Zilliz Cloud는 번거로움이 없으며, Milvus 기반으로 10배 더 빠릅니다.

시작하기
피드백

이 페이지가 도움이 되었나요?