Smallest Token Lengths
Context
A document indexer summarizes short words for a compact report. Given the tokens from one document and a requested count, identify the shortest token lengths. Repeated lengths must remain repeated because each token is a separate document item.
Problem
Write a function that receives a list of document tokens and a requested count. Compute each token's length and select at most the requested number of smallest lengths. Return the selected lengths in nondecreasing order. Each occurrence counts separately: if several tokens have the same length, enough copies must be retained when they rank among the requested smallest values. If the requested count is zero, return an empty list. The input and request are constrained so the request never exceeds the number of tokens.
Examples
limit = 2tokens = ["cat", "elephant", "ox", "rabbit"][2, 3]limit = 3tokens = ["apple", "kiwi", "pear", "plum"][4, 4, 4]limit = 0tokens = ["sun", "moon", "star"][]Constraints
0 <= len(tokens) <= 100001 <= len(tokens[i]) <= 40 for every itokens[i]: alphabet lowercase0 <= limit <= len(tokens)Each token contains lowercase English letters only.- Types:
tokensis str[],limitis int; result is int[]
Function signature
def smallest_token_lengths(tokens: list[str], limit: int) -> list[int]
tokenslist[str]- The document tokens whose character lengths are candidates.
limitint- The maximum number of token lengths to return.
- returns list[int]
- The lengths of the limit shortest tokens, or all available lengths when fewer exist, arranged in nondecreasing order.
Adapted from MBPP problem task_10 (CC BY 4.0). Rewritten, extended and verified by Iksha.
Notes
- The returned list may be empty only when limit is zero or the token list is empty.
- Equal-length tokens retain their multiplicity.