Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 7 additions & 5 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -68,10 +68,12 @@ jobs:
- run: >-
wheel-test/bin/python -c
"import asyncio, importlib.metadata;
from trimwise import ContextSourceResult, ContextTrimResult, Trimmer;
assert importlib.metadata.version('trimwise') == '0.4.1';
sync_result = Trimmer().trim_context(['a'], 1, unit='characters');
async_result = asyncio.run(Trimmer().atrim_context(['a'], 1, unit='characters'));
from trimwise import ContextSource, ContextSourceResult, ContextTrimResult, Trimmer;
assert importlib.metadata.version('trimwise') == '0.5.0';
sources = [ContextSource('a', 'Source: A\n')];
sync_result = Trimmer().trim_context(sources, 11, unit='characters');
async_result = asyncio.run(Trimmer().atrim_context(sources, 11, unit='characters'));
assert isinstance(sync_result, ContextTrimResult);
assert isinstance(sync_result.sources[0], ContextSourceResult);
assert async_result == sync_result"
assert async_result == sync_result;
assert sync_result.text == 'Source: A\na'"
8 changes: 5 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -134,11 +134,11 @@ Keep each responsibility in its existing file. Generated folders such as `dist/`
### Package source

- `src/trimwise/__init__.py`: exposes the supported public API names and nothing else.
- `src/trimwise/composition.py`: reconstructs exact source outputs, local spans, affordable omission
markers, and tiny-budget fallbacks.
- `src/trimwise/composition.py`: reconstructs exact source outputs, local spans, budgeted context
wrappers, affordable omission markers, and tiny-budget fallbacks.
- `src/trimwise/measurement.py`: measures token, word, and character budgets and finds fitting
prefixes.
- `src/trimwise/models.py`: defines public enums, configuration, result values, and semantic
- `src/trimwise/models.py`: defines public enums, inputs, configuration, result values, and semantic
backend errors.
- `src/trimwise/ranking.py`: builds scoring-only section and neighbor context, then implements
structural, BM25, semantic, hybrid, signal, cosine, and MMR ranking calculations.
Expand All @@ -159,6 +159,8 @@ Keep each responsibility in its existing file. Generated folders such as `dist/`
- `tests/test_async_semantic.py`: checks FastEmbed and caller callback precedence, vector validation,
staged failures, model reuse, concurrency, async equivalence, and cancellation behavior.
- `tests/test_context.py`: checks shared-budget validation, selection, composition, counts, and spans.
- `tests/test_context_rendering.py`: checks budgeted source prefixes, separators, rendered counts,
wrapper isolation, and sync/async parity.
- `tests/test_context_semantic.py`: checks context semantic batching, deduplication, and async use.
- `tests/test_docstrings.py`: enforces Python docstrings and verifies that `py.typed` is packaged.
- `tests/test_ranking.py`: checks BM25, centrality, semantic and hybrid fusion, signal scoring,
Expand Down
27 changes: 19 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,23 +77,34 @@ print(result.spans) # Original-input Python-string offsets

### Many sources, one shared limit

Use `trim_context()` when passages from several sources should compete for one evidence budget:
Use `trim_context()` when passages from several sources should compete for one budget. Add
`ContextSource` prefixes when the final rendered labels must fit inside that same limit:

```python
from trimwise import ContextSource, Trimmer

result = Trimmer().trim_context(
[record["text"] for record in records],
[
ContextSource(
text=record["text"],
prefix=f"Source: {record['title']}\nURL: {record['url']}\n",
)
for record in records
],
limit=800,
query="Which recommendations are supported by the reports?",
separator="\n\n",
)

for source in result.sources:
print(records[source.source_index]["url"], source.text)
prompt_ready_context = result.text
assert result.output_count <= result.limit
```

The result keeps one row per input source, including empty excerpts, and the sum of its source
output counts stays within `limit`. Labels, URLs, caller-added headings, separators, instructions,
and answer space are outside that limit. See [Many Sources, One Shared Limit](https://trimwise.readthedocs.io/en/latest/multi-source-context/)
for the complete contract and the difference from `atrim_many()`.
The result keeps one row per input source, including empty excerpts. Prefixes are emitted only for
sources that contribute evidence, and `result.text` contains the fully measured rendering. Your
surrounding instructions and answer space remain outside this limit. Plain string sources still use
the original evidence-only accounting. See [Many Sources, One Shared Limit](https://trimwise.readthedocs.io/en/latest/multi-source-context/)
for both modes and the difference from `atrim_many()`.

Depending on the trimming strategy you want to use, find the corresponding starter code example - [auto](https://trimwise.readthedocs.io/en/latest/strategies/#auto-the-lightweight-default),
[structural](https://trimwise.readthedocs.io/en/latest/strategies/#structural-cover-a-document-without-a-query), [lexical](https://trimwise.readthedocs.io/en/latest/strategies/#lexical-preserve-exact-query-evidence),
Expand Down
10 changes: 9 additions & 1 deletion docs/api-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,13 @@ duplicate contextual passages once per normalized-query callback batch.

::: trimwise.TrimInput

## Shared-context inputs

Use [`ContextSource`][trimwise.ContextSource] when a source needs an output prefix that counts
toward the same shared limit as its retained evidence.

::: trimwise.ContextSource

## Configuration

Use [`TrimConfig`][trimwise.TrimConfig] to configure token encoding, managed embeddings, MMR,
Expand All @@ -41,7 +48,8 @@ measured counts, resolved strategy, and trimming status.
::: trimwise.TrimResult

The context methods return a [`ContextTrimResult`][trimwise.ContextTrimResult] containing one
input-aligned [`ContextSourceResult`][trimwise.ContextSourceResult] per source.
input-aligned [`ContextSourceResult`][trimwise.ContextSourceResult] per source and, for rendered
calls, the complete prompt-ready text.

::: trimwise.ContextTrimResult

Expand Down
33 changes: 20 additions & 13 deletions docs/configuration-and-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,7 @@ Import supported objects directly from `trimwise`:
```python
from trimwise import (
BudgetUnit,
ContextSource,
ContextSourceResult,
ContextTrimResult,
SemanticBackendError,
Expand All @@ -55,6 +56,7 @@ These are the package's documented exports:
| `TrimConfig` | Stores immutable reusable configuration |
| `TrimInput` | Describes one independent request for asynchronous batch trimming |
| `TrimResult` | Reports the excerpt, measurements, resolved strategy, and whether text changed |
| `ContextSource` | Pairs evidence with an optional output-only prefix |
| `ContextSourceResult` | Reports one input-aligned source excerpt and its local spans |
| `ContextTrimResult` | Reports all source excerpts and their shared aggregate measurements |
| `SourceSpan` | Identifies one retained range in the original input string |
Expand Down Expand Up @@ -255,31 +257,34 @@ a boolean.

## `trim_context()` and `atrim_context()`

Use the context methods when many source strings should share one output limit:
Use the context methods when many sources should share one output limit:

```text
trim_context(
sources: Sequence[str],
sources: Sequence[str | ContextSource],
limit: int,
*,
unit: BudgetUnit | str = BudgetUnit.TOKENS,
strategy: Strategy | str = Strategy.AUTO,
query: str | None = None,
token_counter: Callable[[str], int] | None = None,
deduplicate: bool = False,
separator: str | None = None,
) -> ContextTrimResult
```

`atrim_context()` accepts the same arguments and returns the same result type asynchronously.
`sources` must be a sequence of strings, not one bare string or an arbitrary iterable. The result
contains one `ContextSourceResult` per input position. A source may receive an empty excerpt, but
its row and `source_index` remain present.

Source input and output strings are measured independently. Their counts sum to the aggregate
counts, and the aggregate output cannot exceed the shared limit. Caller-added labels, URLs,
instructions, and separators are not counted. See
[Many Sources, One Shared Limit](multi-source-context.md) for prompt assembly and token-counting
guidance.
`sources` must be a sequence of strings or `ContextSource` values, not one bare string or an
arbitrary iterable. `ContextSource(text, prefix="...")` attaches exact output text that is emitted
only if that source contributes evidence. The result contains one `ContextSourceResult` per input
position. A source may receive an empty excerpt, but its row and `source_index` remain present.

Supplying any `ContextSource` or an explicit `separator` returns the complete rendering in
`result.text`. Its aggregate `output_count` measures prefixes, evidence, separators, and omission
text together and cannot exceed the shared limit. Prefixes and separators never affect ranking,
embedding input, or source spans. With plain strings and no separator, `result.text` remains `None`
and the original sum-of-row-counts behavior is preserved. See
[Many Sources, One Shared Limit](multi-source-context.md) for examples and exact counting rules.

`deduplicate=True` is a best-effort embedding option. It sends each exact repeated contextual
passage once during the operation and maps the vector back to every occurrence. It does not remove
Expand Down Expand Up @@ -562,14 +567,16 @@ The context methods return a frozen, slotted `ContextTrimResult`:
| --- | --- |
| `sources` | Input-aligned tuple of `ContextSourceResult` values |
| `input_count` | Sum of all independently measured source inputs |
| `output_count` | Sum of all independently measured source outputs |
| `output_count` | Complete rendered size, or the legacy sum of source outputs |
| `limit` and `unit` | Shared output ceiling and measurement rule |
| `strategy` | Concrete strategy after resolving `auto` |
| `trimmed` | Whether any source output differs from its input |
| `text` | Complete rendered context, or `None` for plain-string calls without a separator |

Each source result contains `source_index`, `text`, `input_count`, `output_count`, `trimmed`, and
local `spans`. Empty or wholly omitted sources keep their row with empty text, zero output count,
and no spans. Keep caller metadata outside the result and reconnect it with `source_index`.
and no spans. Prefixes and separators have no spans. Keep non-rendered caller metadata outside the
result and reconnect it with `source_index`.

## `Strategy`

Expand Down
27 changes: 17 additions & 10 deletions docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -228,22 +228,28 @@ That example gives every source its own 120-token limit. When the sources should
for one allowance, call `trim_context()`:

```python
from trimwise import ContextSource

shared = trimmer.trim_context(
[source["text"] for source in sources],
[
ContextSource(
text=source["text"],
prefix=f"## {source['label']}\n\n",
)
for source in sources
],
limit=300,
query=task,
separator="\n\n",
)

evidence = []
for row in shared.sources:
if row.text:
label = sources[row.source_index]["label"]
evidence.append(f"## {label}\n\n{row.text}")
evidence = shared.text
```

A more relevant source may use more room, and some source rows may be empty. The labels and
separators added above are not part of the 300-token limit. Read
[Many Sources, One Shared Limit](multi-source-context.md) for counts, spans, async use, and the
separators above are emitted only for contributing sources and are included in the 300-token limit.
Instructions around `evidence` and the model's answer still need separate room. Read [Many Sources,
One Shared Limit](multi-source-context.md) for evidence-only mode, counts, spans, async use, and the
difference from `atrim_many()`.

Keep instructions outside the source text passed to Trimwise. The library is designed to reduce
Expand Down Expand Up @@ -415,8 +421,9 @@ not factual truth.
- **Omitting the query:** lexical, semantic, and hybrid strategies require a nonblank query.
- **Treating the limit as a target:** it is a ceiling. A query-aware result may stop early instead
of adding weak evidence.
- **Budgeting only the excerpts:** leave space for labels, separators, instructions, examples, and
the model's answer.
- **Forgetting prompt overhead:** use `ContextSource` when per-source labels and separators must
share the evidence limit, and still reserve room for surrounding instructions, examples, tools,
and the model's answer.
- **Passing instructions as evidence:** trim source material, then assemble it around instructions
that remain unchanged.
- **Expecting a summary:** Trimwise selects and joins original fragments; it does not paraphrase or
Expand Down
37 changes: 23 additions & 14 deletions docs/guarantees-and-limitations.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,14 +28,25 @@ headings, and affordable omission markers have been applied. It returns only whe
result.output_count <= result.limit
```

For `trim_context()` and `atrim_context()`, the ceiling applies to the sum of independently
measured source outputs:
With plain string sources and no explicit separator, context calls preserve their original
row-oriented accounting:

```text
result.output_count = sum(source.output_count for source in result.sources)
result.output_count <= result.limit
```

Passing a `ContextSource` or an explicit separator instead applies the ceiling to the complete
prompt-ready rendering:

```text
measure(result.text) = result.output_count
result.output_count <= result.limit
```

The complete measurement includes prefixes only for contributing sources and separators only
between them. Per-source rows and spans continue to describe the retained source evidence.

The guarantee uses the requested measurement rule:

- Tiktoken or your custom counter for token budgets.
Expand All @@ -52,13 +63,11 @@ for unit, limit in (("tokens", 30), ("words", 20), ("characters", 100)):
assert result.output_count <= limit
```

This guarantee applies to `result.text`, not to the larger prompt around it. Instructions, source
labels, separators, examples, tool definitions, output schemas, and the model's answer all need
their own context space.

The same boundary applies to context results: caller-added labels and separators are not counted.
Separately measured token strings can also tokenize differently after they are joined. Measure the
completed prompt and reserve room when its whole-token ceiling must be exact.
For a single-source `TrimResult`, this guarantee applies to `result.text`, not to the larger prompt
around it. For a context result, prefixes and separators supplied through the context API are
included; formatting added afterward is not. Instructions, examples, tool definitions, output
schemas, and the model's answer still need their own context space. Measure the completed prompt and
reserve room when its whole-token ceiling must be exact.

### Input that already fits is returned exactly

Expand Down Expand Up @@ -384,20 +393,20 @@ Trimwise deliberately leaves several decisions with the application:

| Responsibility | What the caller should do |
| --- | --- |
| Complete prompt budget | Reserve room for labels, instructions, examples, tools, and model output |
| Source identity | Store document IDs, URLs, authors, timestamps, and access controls outside Trimwise results; reconnect context rows with `source_index` |
| Complete prompt budget | Use `ContextSource` for budgeted per-source prefixes; reserve room for surrounding instructions, examples, tools, and model output |
| Source identity | Keep document IDs, authors, timestamps, and access controls in application data; render only the bounded labels you need and reconnect rows with `source_index` |
| Relevance evaluation | Test downstream answers on representative documents and queries |
| Factual verification | Check claims against original sources when accuracy matters |
| Contradiction handling | Preserve and compare conflicting evidence explicitly |
| Embedding operations | Own callback caching, retries, timeouts, rate limits, privacy, and concurrency |
| Security | Treat selected source text as untrusted and enforce tool permissions separately |
| Security | Treat source text and caller-supplied prefixes as untrusted and enforce tool permissions separately |
| Tokenizer alignment | Supply a custom counter when the default encoding does not match the target model |

`TrimResult.spans` exposes ordered Python-string ranges for retained source text, but it does not
include source IDs, scores, or embeddings. Starts are inclusive, ends are exclusive, and adjacent
ranges are merged. Overlapping ranges are never returned because candidates do not overlap.
Generated omission markers and separators have no span. Keep document identity alongside each
input before trimming several sources.
Generated omission markers and caller-supplied prefixes and separators have no span. Keep document
identity alongside each input before trimming several sources.

For a context result, each source row's spans index only the original string at its `source_index`.

Expand Down
Loading
Loading