語言

CodeGenTokenizer.GetIndexByTokenCountFromEnd 方法

定義

多載

名稱 Description
GetIndexByTokenCountFromEnd(ReadOnlySpan<Char>, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)

找出從文字末尾最大編碼容量的索引,且不超過標記限制。

GetIndexByTokenCountFromEnd(String, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)

找出從文字末尾最大編碼容量的索引,且不超過標記限制。

GetIndexByTokenCountFromEnd(ReadOnlySpan<Char>, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)

來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs

找出從文字末尾最大編碼容量的索引,且不超過標記限制。

public int GetIndexByTokenCountFromEnd(ReadOnlySpan<char> text, int maxTokenCount, bool addPrefixSpace, bool addBeginningOfSentence, bool addEndOfSentence, out string? normalizedText, out int tokenCount, bool considerPreTokenization = true, bool considerNormalization = true);
override this.GetIndexByTokenCountFromEnd : ReadOnlySpan<char> * int * bool * bool * bool * string * int * bool * bool -> int
Public Function GetIndexByTokenCountFromEnd (text As ReadOnlySpan(Of Char), maxTokenCount As Integer, addPrefixSpace As Boolean, addBeginningOfSentence As Boolean, addEndOfSentence As Boolean, ByRef normalizedText As String, ByRef tokenCount As Integer, Optional considerPreTokenization As Boolean = true, Optional considerNormalization As Boolean = true) As Integer

參數

text
ReadOnlySpan<Char>

要編碼的文字。

maxTokenCount
Int32

限制編碼容量的最大令牌數。

addPrefixSpace
Boolean

在編碼文字前,請指示是否要加上導引空格。

addBeginningOfSentence
Boolean

指示是否要在編碼中包含句子標記的開頭。

addEndOfSentence
Boolean

指示是否要在編碼中包含句尾的標記。

normalizedText
String

若啟用分詞器的正規化,輸入文字將以正規化形式呈現;否則,該數字將為零。

tokenCount
Int32

代幣數量應該比最大代幣數量還要少。

considerPreTokenization
Boolean

說明是否在分詞化前考慮預先分詞化。

considerNormalization
Boolean

說明是否在分詞前考慮正規化。

傳回

處理後的文字中最大編碼容量的起始索引,且不超過標記限制。 它代表第一個字元的索引。 若無標記匹配,結果為 長度 normalizedText;反之,若所有標記匹配,結果為 0。

備註

如果整段文字都能在標記限制內編碼,回傳的索引將為 0。

適用於

GetIndexByTokenCountFromEnd(String, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)

來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs
來源:
CodeGenTokenizer.cs

找出從文字末尾最大編碼容量的索引,且不超過標記限制。

public int GetIndexByTokenCountFromEnd(string text, int maxTokenCount, bool addPrefixSpace, bool addBeginningOfSentence, bool addEndOfSentence, out string? normalizedText, out int tokenCount, bool considerPreTokenization = true, bool considerNormalization = true);
override this.GetIndexByTokenCountFromEnd : string * int * bool * bool * bool * string * int * bool * bool -> int
Public Function GetIndexByTokenCountFromEnd (text As String, maxTokenCount As Integer, addPrefixSpace As Boolean, addBeginningOfSentence As Boolean, addEndOfSentence As Boolean, ByRef normalizedText As String, ByRef tokenCount As Integer, Optional considerPreTokenization As Boolean = true, Optional considerNormalization As Boolean = true) As Integer

參數

text
String

要編碼的文字。

maxTokenCount
Int32

限制編碼容量的最大令牌數。

addPrefixSpace
Boolean

在編碼文字前,請指示是否要加上導引空格。

addBeginningOfSentence
Boolean

指示是否要在編碼中包含句子標記的開頭。

addEndOfSentence
Boolean

指示是否要在編碼中包含句尾的標記。

normalizedText
String

若啟用分詞器的正規化,輸入文字將以正規化形式呈現;否則,該數字將為零。

tokenCount
Int32

代幣數量應該比最大代幣數量還要少。

considerPreTokenization
Boolean

說明是否在分詞化前考慮預先分詞化。

considerNormalization
Boolean

說明是否在分詞前考慮正規化。

傳回

處理後的文字中最大編碼容量的起始索引,且不超過標記限制。 它代表第一個字元的索引。 若無標記匹配,結果為文字長度,若啟用正規化則為 ; normalizedText 反之,若所有標記皆符合,結果為 0。

備註

如果整段文字都能在標記限制內編碼,回傳的索引將為 0。

適用於