語言

CodeGenTokenizer.EncodeToIds 方法

定義

多載

名稱 Description
EncodeToIds(String, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)

將輸入文字編碼為標記 Id,最多可達標記數量。

EncodeToIds(String, Boolean, Boolean, Boolean, Boolean, Boolean)

將輸入文字編碼為 Ids。

EncodeToIds(ReadOnlySpan<Char>, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)

將輸入文字編碼為標記 Id,最多可達標記數量。

EncodeToIds(String, ReadOnlySpan<Char>, EncodeSettings)

將輸入文字編碼為標記 ID。

EncodeToIds(ReadOnlySpan<Char>, Boolean, Boolean, Boolean, Boolean, Boolean)

將輸入文字編碼為 Ids。

EncodeToIds(String, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)

來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs

將輸入文字編碼為標記 Id,最多可達標記數量。

public System.Collections.Generic.IReadOnlyList<int> EncodeToIds(string text, int maxTokenCount, bool addPrefixSpace, bool addBeginningOfSentence, bool addEndOfSentence, out string? normalizedText, out int charsConsumed, bool considerPreTokenization = true, bool considerNormalization = true);
override this.EncodeToIds : string * int * bool * bool * bool * string * int * bool * bool -> System.Collections.Generic.IReadOnlyList<int>
Public Function EncodeToIds (text As String, maxTokenCount As Integer, addPrefixSpace As Boolean, addBeginningOfSentence As Boolean, addEndOfSentence As Boolean, ByRef normalizedText As String, ByRef charsConsumed As Integer, Optional considerPreTokenization As Boolean = true, Optional considerNormalization As Boolean = true) As IReadOnlyList(Of Integer)

參數

text
String

要編碼的文字。

maxTokenCount
Int32

最多要編碼的標記數量。

addPrefixSpace
Boolean

在編碼文字前,請指示是否要加上導引空格。

addBeginningOfSentence
Boolean

指示是否要在編碼中包含句子標記的開頭。

addEndOfSentence
Boolean

指示是否要在編碼中包含句尾的標記。

normalizedText
String

若啟用分詞器的正規化,輸入文字將以正規化形式呈現;否則,該數字將為零。

charsConsumed
Int32

包含最大編碼標記的文字長度。

considerPreTokenization
Boolean

說明是否在分詞化前考慮預先分詞化。

considerNormalization
Boolean

說明是否在分詞前考慮正規化。

傳回

編碼的 ID 清單。

適用於

EncodeToIds(String, Boolean, Boolean, Boolean, Boolean, Boolean)

來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs

將輸入文字編碼為 Ids。

public System.Collections.Generic.IReadOnlyList<int> EncodeToIds(string text, bool addPrefixSpace, bool addBeginningOfSentence, bool addEndOfSentence, bool considerPreTokenization = true, bool considerNormalization = true);
override this.EncodeToIds : string * bool * bool * bool * bool * bool -> System.Collections.Generic.IReadOnlyList<int>
Public Function EncodeToIds (text As String, addPrefixSpace As Boolean, addBeginningOfSentence As Boolean, addEndOfSentence As Boolean, Optional considerPreTokenization As Boolean = true, Optional considerNormalization As Boolean = true) As IReadOnlyList(Of Integer)

參數

text
String

要編碼的文字。

addPrefixSpace
Boolean

在編碼文字前,請指示是否要加上導引空格。

addBeginningOfSentence
Boolean

指示是否要在編碼中包含句子標記的開頭。

addEndOfSentence
Boolean

指示是否要在編碼中包含句尾的標記。

considerPreTokenization
Boolean

說明是否在分詞化前考慮預先分詞化。

considerNormalization
Boolean

說明是否在分詞前考慮正規化。

傳回

編碼的 ID 清單。

適用於

EncodeToIds(ReadOnlySpan<Char>, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)

來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs

將輸入文字編碼為標記 Id,最多可達標記數量。

public System.Collections.Generic.IReadOnlyList<int> EncodeToIds(ReadOnlySpan<char> text, int maxTokenCount, bool addPrefixSpace, bool addBeginningOfSentence, bool addEndOfSentence, out string? normalizedText, out int charsConsumed, bool considerPreTokenization = true, bool considerNormalization = true);
override this.EncodeToIds : ReadOnlySpan<char> * int * bool * bool * bool * string * int * bool * bool -> System.Collections.Generic.IReadOnlyList<int>
Public Function EncodeToIds (text As ReadOnlySpan(Of Char), maxTokenCount As Integer, addPrefixSpace As Boolean, addBeginningOfSentence As Boolean, addEndOfSentence As Boolean, ByRef normalizedText As String, ByRef charsConsumed As Integer, Optional considerPreTokenization As Boolean = true, Optional considerNormalization As Boolean = true) As IReadOnlyList(Of Integer)

參數

text
ReadOnlySpan<Char>

要編碼的文字。

maxTokenCount
Int32

最多要編碼的標記數量。

addPrefixSpace
Boolean

在編碼文字前,請指示是否要加上導引空格。

addBeginningOfSentence
Boolean

指示是否要在編碼中包含句子標記的開頭。

addEndOfSentence
Boolean

指示是否要在編碼中包含句尾的標記。

normalizedText
String

若啟用分詞器的正規化,輸入文字將以正規化形式呈現;否則,該數字將為零。

charsConsumed
Int32

包含最大編碼標記的文字長度。

considerPreTokenization
Boolean

說明是否在分詞化前考慮預先分詞化。

considerNormalization
Boolean

說明是否在分詞前考慮正規化。

傳回

編碼的 ID 清單。

適用於

EncodeToIds(String, ReadOnlySpan<Char>, EncodeSettings)

來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs

將輸入文字編碼為標記 ID。

protected override Microsoft.ML.Tokenizers.EncodeResults<int> EncodeToIds(string? text, ReadOnlySpan<char> textSpan, Microsoft.ML.Tokenizers.EncodeSettings settings);
override this.EncodeToIds : string * ReadOnlySpan<char> * Microsoft.ML.Tokenizers.EncodeSettings -> Microsoft.ML.Tokenizers.EncodeResults<int>
Protected Overrides Function EncodeToIds (text As String, textSpan As ReadOnlySpan(Of Char), settings As EncodeSettings) As EncodeResults(Of Integer)

參數

text
String

要編碼的文字。

textSpan
ReadOnlySpan<Char>

如果 為 textnull,則會使用的文字範圍。

settings
EncodeSettings

編碼文字的設定。

傳回

包含編碼 ID 列表的編碼結果。

適用於

EncodeToIds(ReadOnlySpan<Char>, Boolean, Boolean, Boolean, Boolean, Boolean)

來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs

將輸入文字編碼為 Ids。

public System.Collections.Generic.IReadOnlyList<int> EncodeToIds(ReadOnlySpan<char> text, bool addPrefixSpace, bool addBeginningOfSentence, bool addEndOfSentence, bool considerPreTokenization = true, bool considerNormalization = true);
override this.EncodeToIds : ReadOnlySpan<char> * bool * bool * bool * bool * bool -> System.Collections.Generic.IReadOnlyList<int>
Public Function EncodeToIds (text As ReadOnlySpan(Of Char), addPrefixSpace As Boolean, addBeginningOfSentence As Boolean, addEndOfSentence As Boolean, Optional considerPreTokenization As Boolean = true, Optional considerNormalization As Boolean = true) As IReadOnlyList(Of Integer)

參數

text
ReadOnlySpan<Char>

要編碼的文字。

addPrefixSpace
Boolean

在編碼文字前,請指示是否要加上導引空格。

addBeginningOfSentence
Boolean

指示是否要在編碼中包含句子標記的開頭。

addEndOfSentence
Boolean

指示是否要在編碼中包含句尾的標記。

considerPreTokenization
Boolean

說明是否在分詞化前考慮預先分詞化。

considerNormalization
Boolean

說明是否在分詞前考慮正規化。

傳回

編碼的 ID 清單。

適用於