CodeGenTokenizer.GetIndexByTokenCountFromEnd 方法
定義
重要
部分資訊涉及發行前產品,在發行之前可能會有大幅修改。 Microsoft 對此處提供的資訊,不做任何明確或隱含的瑕疵擔保。
多載
| 名稱 | Description |
|---|---|
| GetIndexByTokenCountFromEnd(ReadOnlySpan<Char>, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean) |
找出從文字末尾最大編碼容量的索引,且不超過標記限制。 |
| GetIndexByTokenCountFromEnd(String, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean) |
找出從文字末尾最大編碼容量的索引,且不超過標記限制。 |
GetIndexByTokenCountFromEnd(ReadOnlySpan<Char>, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)
找出從文字末尾最大編碼容量的索引,且不超過標記限制。
public int GetIndexByTokenCountFromEnd(ReadOnlySpan<char> text, int maxTokenCount, bool addPrefixSpace, bool addBeginningOfSentence, bool addEndOfSentence, out string? normalizedText, out int tokenCount, bool considerPreTokenization = true, bool considerNormalization = true);
override this.GetIndexByTokenCountFromEnd : ReadOnlySpan<char> * int * bool * bool * bool * string * int * bool * bool -> int
Public Function GetIndexByTokenCountFromEnd (text As ReadOnlySpan(Of Char), maxTokenCount As Integer, addPrefixSpace As Boolean, addBeginningOfSentence As Boolean, addEndOfSentence As Boolean, ByRef normalizedText As String, ByRef tokenCount As Integer, Optional considerPreTokenization As Boolean = true, Optional considerNormalization As Boolean = true) As Integer
參數
- text
- ReadOnlySpan<Char>
要編碼的文字。
- maxTokenCount
- Int32
限制編碼容量的最大令牌數。
- addPrefixSpace
- Boolean
在編碼文字前,請指示是否要加上導引空格。
- addBeginningOfSentence
- Boolean
指示是否要在編碼中包含句子標記的開頭。
- addEndOfSentence
- Boolean
指示是否要在編碼中包含句尾的標記。
- normalizedText
- String
若啟用分詞器的正規化,輸入文字將以正規化形式呈現;否則,該數字將為零。
- tokenCount
- Int32
代幣數量應該比最大代幣數量還要少。
- considerPreTokenization
- Boolean
說明是否在分詞化前考慮預先分詞化。
- considerNormalization
- Boolean
說明是否在分詞前考慮正規化。
傳回
處理後的文字中最大編碼容量的起始索引,且不超過標記限制。
它代表第一個字元的索引。 若無標記匹配,結果為 長度 normalizedText;反之,若所有標記匹配,結果為 0。
備註
如果整段文字都能在標記限制內編碼,回傳的索引將為 0。
適用於
GetIndexByTokenCountFromEnd(String, Int32, Boolean, Boolean, Boolean, String, Int32, Boolean, Boolean)
找出從文字末尾最大編碼容量的索引,且不超過標記限制。
public int GetIndexByTokenCountFromEnd(string text, int maxTokenCount, bool addPrefixSpace, bool addBeginningOfSentence, bool addEndOfSentence, out string? normalizedText, out int tokenCount, bool considerPreTokenization = true, bool considerNormalization = true);
override this.GetIndexByTokenCountFromEnd : string * int * bool * bool * bool * string * int * bool * bool -> int
Public Function GetIndexByTokenCountFromEnd (text As String, maxTokenCount As Integer, addPrefixSpace As Boolean, addBeginningOfSentence As Boolean, addEndOfSentence As Boolean, ByRef normalizedText As String, ByRef tokenCount As Integer, Optional considerPreTokenization As Boolean = true, Optional considerNormalization As Boolean = true) As Integer
參數
- text
- String
要編碼的文字。
- maxTokenCount
- Int32
限制編碼容量的最大令牌數。
- addPrefixSpace
- Boolean
在編碼文字前,請指示是否要加上導引空格。
- addBeginningOfSentence
- Boolean
指示是否要在編碼中包含句子標記的開頭。
- addEndOfSentence
- Boolean
指示是否要在編碼中包含句尾的標記。
- normalizedText
- String
若啟用分詞器的正規化,輸入文字將以正規化形式呈現;否則,該數字將為零。
- tokenCount
- Int32
代幣數量應該比最大代幣數量還要少。
- considerPreTokenization
- Boolean
說明是否在分詞化前考慮預先分詞化。
- considerNormalization
- Boolean
說明是否在分詞前考慮正規化。
傳回
處理後的文字中最大編碼容量的起始索引,且不超過標記限制。
它代表第一個字元的索引。 若無標記匹配,結果為文字長度,若啟用正規化則為 ; normalizedText 反之,若所有標記皆符合,結果為 0。
備註
如果整段文字都能在標記限制內編碼,回傳的索引將為 0。