API Reference¶
extract_pii¶
Extract PII entities from text with intelligent entity merging.
Uses token classification models to detect personally identifiable information including names, emails, phone numbers, addresses, and other HIPAA-protected identifiers.
The smart merging feature uses regex patterns to identify semantic units (dates, SSN, phone numbers, etc.) and merges fragmented model predictions into complete entities with dominant label selection.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text | str | bytes | bytearray | memoryview | Input text to analyze | required |
model_name | str | PII detection model (registry key or HuggingFace ID). When the default is used and | _DEFAULT_EN_MODEL |
confidence_threshold | float | Minimum confidence score (0-1) | 0.5 |
config | Optional[OpenMedConfig] | Optional configuration override | None |
use_smart_merging | bool | Enable regex-based semantic unit merging (recommended) | True |
lang | str | ISO 639-1 language code (en, fr, de, it, es, nl, hi, te, pt, ar, ja, tr). Controls which default model and regex patterns are used. Mixed Latin/Devanagari or Latin/Telugu notes automatically use script-aware India clinical routing for | 'en' |
normalize_accents | Optional[bool] | Strip diacritical marks before model inference so that models trained on accent-free text still detect accented names. Entity spans in the result reference the original (accented) text. | None |
preserve_whitespace | bool | Preserve leading and trailing source whitespace so returned entity offsets refer to the exact input string. | False |
loader | Optional['ModelLoader'] | Optional shared model loader to reuse warmed pipelines. | None |
batch_size | Optional[int] | Optional backend inference batch size. | None |
num_workers | Optional[int] | Optional backend inference worker count. | None |
custom_recognizer | Any | Optional deny-list/allow-list recognizer config, | None |
abdm | Optional[bool] | Enable the India ABDM identifier bundle. | None |
code_mixed | bool | Enable the explicit English/Hinglish route. This preserves English model detection while adding Roman-Hindi context patterns. | False |
token_language_tags | Optional[Sequence[Any]] | Required with | None |
lid_model | Optional['TokenLIDHook'] | Optional user-supplied token language-ID hook used only when | None |
transliterated_name_config | Any | Optional configuration for the conservative Latin-script Indian given/family-name allow/deny bridge. | None |
cache_results | bool | Whether to cache this result in the in-process LRU cache. Cached results may contain PHI, but are never saved to disk. | False |
max_cache_entries | int | Maximum number of cached results. | 128 |
budget | Optional[RequestBudget] | Optional per-request wall-time and input-character budget. | None |
Returns:
| Type | Description |
|---|---|
PredictionResult | PredictionResult with detected PII entities |
Raises:
| Type | Description |
|---|---|
InputError | If text or request options are malformed or unsupported. |
CapabilityError | If the selected model or optional runtime is unavailable. |
BudgetExceededError | If the request exceeds its configured budget. |
InternalError | If inference violates a result or span invariant. |
Example
from unittest.mock import patch from openmed.core.pii import extract_pii from openmed.processing.outputs import EntityPrediction, PredictionResult fake_result = PredictionResult( ... text="Patient Casey Example called.", ... entities=[ ... EntityPrediction( ... text="Casey Example", ... label="NAME", ... confidence=0.98, ... start=8, ... end=21, ... ) ... ], ... model_name="fixture-pii-model", ... timestamp="2026-01-01T00:00:00", ... ) with patch("openmed.analyze_text", return_value=fake_result): ... result = extract_pii( ... "Patient Casey Example called.", ... model_name="fixture-pii-model", ... use_smart_merging=False, ... ) next((entity.text, entity.label) for entity in result.entities) ('Casey Example', 'NAME')
deidentify¶
De-identify text by detecting and redacting PII with intelligent merging.
Implements multiple de-identification strategies for HIPAA compliance:
- mask: Replace with placeholders like [NAME], [EMAIL], etc.
- aadhaar_mask: Render valid Aadhaar values as
XXXX XXXX NNNN; use ordinary placeholders for every other entity - remove: Remove PII text entirely (empty string)
- replace: Replace with fake but realistic data
- hash: Replace with consistent hashed values for entity linking
- format_preserve: Replace structured identifiers with synthetic values that keep shape and separators, masking unsupported labels
- shift_dates: Shift dates by random offset while preserving intervals
Smart merging uses regex patterns to merge fragmented entities (e.g., dates split into '01' and '/15/1970' are merged into complete '01/15/1970').
Code-mixed mode is explicit and offset driven. With code_mixed=True and per-token language tags, the English NER path remains active while a separate Roman-script Hindi pattern bank detects cues such as naam, umar, pata, mobile, and janm. The combined spans pass through the normal entity merger and final safety sweep before redaction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text | str | bytes | bytearray | memoryview | Input text to de-identify | required |
method | DeidentificationMethod | De-identification method (mask, aadhaar_mask, remove, replace, hash, shift_dates, format_preserve) | 'mask' |
model_name | str | PII detection model | _DEFAULT_EN_MODEL |
confidence_threshold | float | Minimum confidence for redaction (default 0.7 for safety) | 0.7 |
keep_year | bool | For dates, keep the year unchanged | False |
shift_dates | Optional[bool] | Deprecated alias for | None |
date_shift_days | Optional[int] | Specific number of days to shift when | None |
patient_key | Optional[str | bytes] | Optional stable patient identifier used only to derive a deterministic HMAC date-shift offset. Raw keys are not logged, persisted, or returned. | None |
date_shift_max_days | Optional[int] | Maximum absolute offset for random, seeded, or patient-keyed date shifting. Defaults to 365 when | None |
date_shift_secret | Optional[str | bytes] | Required HMAC key material for patient-keyed offsets. Reuse the same value across sessions to keep offsets stable. | None |
keep_mapping | bool | Keep mapping for re-identification | False |
config | Optional[OpenMedConfig] | Optional configuration override | None |
use_smart_merging | bool | Enable regex-based semantic unit merging (recommended) | True |
use_safety_sweep | bool | Run a deterministic structured-identifier sweep after model detection and before redaction. | True |
lang | str | ISO 639-1 language code (en, fr, de, it, es, nl, hi, te, pt, ar, ja, tr). Controls model selection, regex patterns, and fake data for replacement. Mixed Latin/Devanagari or Latin/Telugu notes automatically use the script-aware India clinical route for | 'en' |
normalize_accents | Optional[bool] | Strip diacritical marks before model inference. | None |
loader | Optional['ModelLoader'] | Optional shared model loader to reuse warmed pipelines. | None |
consistent | bool | When | False |
seed | Optional[int] | Optional request-scoped integer seed for cross-run reproducibility of replacements and automatic date shifting. Implies | None |
locale | Optional[str] | Faker locale override ( | None |
surrogate_vault | Optional['SurrogateVault'] | Optional cross-document surrogate vault. When provided with | None |
policy | Optional[str] | Optional policy profile name controlling arbitration, action selection, mandatory safety sweep behavior, and reversible mapping. | None |
calibration_thresholds_path | Optional[str | Path] | Optional thresholds.json artifact path or artifact directory. When provided, per-label calibrated thresholds filter model detections and appear in audit output. | None |
custom_recognizer | Any | Optional deny-list/allow-list recognizer config, | None |
abdm | Optional[bool] | Enable the India ABDM recognizer bundle. | None |
code_mixed | bool | Enable the explicit English/Hinglish de-identification path. | False |
token_language_tags | Optional[Sequence[Any]] | Required with | None |
lid_model | Optional['TokenLIDHook'] | Optional user-supplied token language-ID hook used only when code-mixed tags are inferred. | None |
transliterated_name_config | Any | Optional configuration for the Latin-script Indian name allow/deny bridge. The default bridge is conservative and can be replaced or extended by configuration. | None |
audit | bool | Return an AuditReport instead of the DeidentificationResult. Fresh calls use separate private HMAC keys. For stable hashes across runs, use a Pipeline with an explicit private hmac_secret. | False |
cache_results | bool | Whether to cache this result in the in-process LRU cache. Cached results may contain PHI, but are never saved to disk. | False |
max_cache_entries | int | Maximum number of cached results. | 128 |
budget | Optional[RequestBudget] | Optional per-request wall-time and input-character budget. | None |
Returns:
| Type | Description |
|---|---|
DeidentificationResult | 'AuditReport' | DeidentificationResult with original and de-identified text, or |
DeidentificationResult | 'AuditReport' | AuditReport when |
Raises:
| Type | Description |
|---|---|
InputError | If text or de-identification options are malformed. |
PolicyError | If the selected privacy policy is invalid or disallows the requested operation. |
CapabilityError | If the selected model or optional runtime is unavailable. |
BudgetExceededError | If the request exceeds its configured budget. |
InternalError | If inference or redaction violates an internal invariant. |
Example
from datetime import datetime from types import SimpleNamespace from unittest.mock import patch from openmed.core.pii import ( ... DeidentificationResult, ... PIIEntity, ... deidentify, ... ) fixture = DeidentificationResult( ... original_text="Patient Casey Example", ... deidentified_text="Patient [NAME]", ... pii_entities=[ ... PIIEntity( ... text="Casey Example", ... label="NAME", ... start=8, ... end=21, ... confidence=0.98, ... redacted_text="[NAME]", ... ) ... ], ... method="mask", ... timestamp=datetime(2026, 1, 1, 0, 0, 0), ... mapping={"[NAME]": "Casey Example"}, ... ) with patch("openmed.core.pipeline.Pipeline") as pipeline_cls: ... pipeline_cls.return_value.run.return_value = SimpleNamespace( ... deidentification_result=fixture ... ) ... result = deidentify( ... "Patient Casey Example", ... method="mask", ... keep_mapping=True, ... ) result.deidentified_text 'Patient [NAME]' result.mapping
reidentify¶
Re-identify text using stored mapping.
Restores original PII from de-identified text using the mapping created during de-identification. Only works if keep_mapping=True was used.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
deidentified_text | str | De-identified text | required |
mapping | Mapping[str, str] | Mapping from redacted to original text, optionally including occurrence-aware entries emitted for colliding replacement values. | required |
Returns:
| Type | Description |
|---|---|
str | Re-identified text with original PII restored |
Raises:
| Type | Description |
|---|---|
InputError | If the text or mapping does not use the documented string types. |
Example
from openmed.core.pii import reidentify reidentify( ... "Patient [NAME] has record [ID]", ... {"[NAME]": "Casey Example", "[ID]": "MRN-0001"}, ... ) 'Patient Casey Example has record MRN-0001'
Note
Only works if keep_mapping=True was used during de-identification. Requires proper authorization and audit logging in production.
Structured errors¶
The complete hierarchy, compatibility guarantees, and REST/MCP mappings are documented in Structured public errors.
OpenMedError¶
Bases: Exception
Base class for expected failures on the OpenMed public API.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
message | str | Actionable, PHI-free explanation of the failure. | required |
code | Optional[str] | Optional stable leaf code owned by OpenMed. Callers should not invent codes; this hook supports existing specialized errors. | None |
details | Optional[Mapping[str, Any]] | Optional PHI-free structured context. | None |
Attributes:
| Name | Type | Description |
|---|---|---|
code | str | Stable machine-readable error code. |
message | Actionable, PHI-free human-readable message. | |
details | dict[str, Any] | PHI-free structured context. |
__str__() ¶
Return the PHI-free human-readable message.
to_dict(*, include_details=True) ¶
Return a JSON-ready error object.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
include_details | bool | Include the structured details mapping. Service adapters use | True |
Returns:
| Type | Description |
|---|---|
dict[str, Any] | A mapping with stable |
InputError¶
Bases: OpenMedError, ValueError, TypeError
Caller input is malformed, conflicting, or unsupported.
ValueError and TypeError remain bases so existing handlers keep catching value and type validation failures after adopting the taxonomy.
ConfigurationError¶
Bases: OpenMedError, ValueError, TypeError, KeyError
Configuration is missing, unknown, or inconsistent.
The legacy ValueError, TypeError, and KeyError bases preserve compatibility with configuration registries and validators.
CapabilityError¶
MissingExtraError¶
Bases: CapabilityError
An optional package or OpenMed extra is not installed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
message | str | Actionable message containing an installation instruction. | required |
package | Optional[str] | Missing distribution name, if known. | None |
feature | Optional[str] | Feature that requires the package, if known. | None |
extra | Optional[str] | OpenMed extra that provides the package, if known. | None |
details | Optional[Mapping[str, Any]] | Additional PHI-free structured context. | None |
ModelLoadError¶
Bases: CapabilityError, ValueError
A model, tokenizer, or inference backend could not be loaded.
ValueError is retained in addition to ImportError because older model-loading paths used both builtin families.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
message | str | Actionable, PHI-free load failure message. | required |
model_name | Optional[str] | Non-sensitive model identifier or local path, if known. | None |
details | Optional[Mapping[str, Any]] | Additional PHI-free structured context. | None |
PolicyError¶
Bases: OpenMedError, ValueError, TypeError
A request violates or misconfigures a privacy policy constraint.
BudgetExceededError¶
Bases: OpenMedError, RuntimeError
A request exceeded a configured size, time, or resource budget.
This class accepts both the generic taxonomy constructor and the historical request-budget fields used by :mod:openmed.core.budget.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
message | Optional[str] | Optional actionable message. When omitted, one is built from | None |
kind | Optional[str] | Budget dimension such as | None |
limit | Optional[float] | Configured limit. | None |
observed | Optional[float] | Observed value that exceeded the limit. | None |
checkpoint | Optional[str] | Safe pipeline checkpoint where the limit was observed. | None |
details | Optional[Mapping[str, Any]] | Additional PHI-free structured context. | None |
InternalError¶
Bases: OpenMedError, RuntimeError
An internal invariant failed and the request cannot safely continue.
InferenceError¶
redact_detail¶
Return a stable descriptor for untrusted text without exposing it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value | Any | Value to describe. It is converted to text only for hashing and is never included verbatim in the returned descriptor. | required |
Returns:
| Type | Description |
|---|---|
str | A descriptor containing only the UTF-8 byte length and SHA-256 digest. |
analyze_text¶
Run a token-classification model on text and format the predictions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text | str | Clinical or biomedical text to analyse. | required |
model_name | str | Registry key, fully-qualified Hugging Face model id, or local model path. | 'disease_detection_superclinical' |
model_id | Optional[str] | Alias for | None |
config | Optional[OpenMedConfig] | Optional :class: | None |
loader | Optional[ModelLoader] | Reuse an existing :class: | None |
aggregation_strategy | Optional[str] | Hugging Face aggregation strategy ( | 'simple' |
output_format | str |
| 'dict' |
include_confidence | bool | Whether to include confidence scores in formatted output. | True |
confidence_threshold | Optional[float] | Minimum confidence for entities. | 0.0 |
group_entities | bool | Merge adjacent entities of the same label in the formatted output. | False |
formatter_kwargs | Optional[Dict[str, Any]] | Extra keyword arguments forwarded to :func: | None |
metadata | Optional[Dict[str, Any]] | Optional metadata to attach to the result. | None |
use_fast_tokenizer | bool | Prefer fast tokenizers when available. | True |
sentence_detection | bool | Enable sentence detection (default: True). The engine is selected by | True |
sentence_language | str | Language hint for the sentence detector. | 'en' |
sentence_clean | bool | Whether to enable the sentence detector's cleaning heuristics. | False |
sentence_segmenter | Optional[Any] | Optional preconstructed segmenter object to reuse. It cannot be combined with | None |
sentence_backend | Literal['auto', 'yasbd'] | Sentence segmentation engine to use. It can be | 'auto' |
assert_context | bool | Attach deterministic negation, uncertainty, experiencer, and temporality labels to each entity under | False |
cache_results | bool | Whether to cache this result in the in-process LRU cache. Cached results may contain PHI, but are never saved to disk. | False |
max_cache_entries | int | Maximum number of cached results. | 128 |
**pipeline_kwargs | Any | Additional arguments passed to :meth: | {} |
Returns:
| Type | Description |
|---|---|
Union[AnalyzeResult, str, List[Dict[str, Any]]] | Analyze result for |
Union[AnalyzeResult, str, List[Dict[str, Any]]] | format. |
Example
class FixtureLoader: ... config = None ... ... def create_pipeline(self, model_name, kwargs): ... def pipeline(text, call_kwargs): ... return [ ... { ... "entity_group": "CONDITION", ... "score": 0.99, ... "start": 11, ... "end": 17, ... "word": "asthma", ... } ... ] ... ... return pipeline ... ... def get_max_sequence_length(self, model_name, tokenizer=None): ... return 128 result = analyze_text( ... "History of asthma.", ... model_name="fixture-ner-model", ... loader=FixtureLoader(), ... sentence_detection=False, ... ) next((entity.text, entity.label) for entity in result.entities) ('asthma', 'CONDITION')
list_models¶
Return available OpenMed model identifiers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
include_registry | bool | Include entries from the bundled registry in addition to entries in the committed manifest. | True |
include_remote | bool | Retained for compatibility; no live discovery is performed. | True |
config | Optional[OpenMedConfig] | Optional custom configuration for model discovery. | None |
BatchProcessor¶
Process multiple texts efficiently with progress tracking.
Example usage
from openmed import BatchProcessor, OpenMedConfig processor = BatchProcessor(model_name="disease_detection_superclinical") texts = ["Patient has diabetes.", "No significant findings."] result = processor.process_texts(texts) print(result.summary())
__init__(model_name='disease_detection_superclinical', operation='analyze_text', batch_size=8, config=None, loader=None, aggregation_strategy='simple', confidence_threshold=None, group_entities=False, continue_on_error=True, checkpoint_interval=_DEFAULT_CHECKPOINT_INTERVAL, _atomic_write_hook=None, budget=None, **analyze_kwargs) ¶
Initialize batch processor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_name | str | Model registry key or HuggingFace identifier. | 'disease_detection_superclinical' |
operation | BatchOperation | Which function to call per item: | 'analyze_text' |
batch_size | int | Number of documents to process together per batch. | 8 |
config | Optional[Any] | Optional OpenMedConfig instance. | None |
loader | Optional[Any] | Optional ModelLoader instance to reuse. | None |
aggregation_strategy | Optional[str] | HuggingFace aggregation strategy ( | 'simple' |
confidence_threshold | Optional[float] | Minimum confidence for entities. When not provided, defaults match the selected operation: | None |
group_entities | bool | Whether to group adjacent entities ( | False |
continue_on_error | bool | Continue processing on individual item errors. | True |
checkpoint_interval | int | Maximum number of items processed between durable checkpoints. | _DEFAULT_CHECKPOINT_INTERVAL |
budget | Optional[Any] | Optional per-request resource budget applied independently to each | None |
**analyze_kwargs | Any | Additional arguments passed to the selected function. | {} |
iter_process(texts, ids=None, *, on_progress=None) ¶
Process texts as an iterator, yielding results one at a time.
This is useful for streaming results or processing very large batches where you don't want to hold all results in memory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
texts | Sequence[str] | Sequence of texts to analyze. | required |
ids | Optional[Sequence[str]] | Optional identifiers for each text. | None |
on_progress | Optional[BatchProgressCallback] | Optional PHI-safe callback that receives a BatchProgress record after each completed item. | None |
Yields:
| Type | Description |
|---|---|
BatchItemResult | BatchItemResult for each processed text. |
process_directory(directory, pattern='*.txt', recursive=False, encoding='utf-8', progress_callback=None, *, on_progress=None, output_path=None, checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, output_format='json') ¶
Process all matching files in a directory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
directory | Union[str, Path] | Directory path. | required |
pattern | str | Glob pattern for file matching. | '*.txt' |
recursive | bool | Whether to search recursively. | False |
encoding | str | File encoding. | 'utf-8' |
progress_callback | Optional[ProgressCallback] | Optional callback for progress updates. | None |
on_progress | Optional[BatchProgressCallback] | Optional PHI-safe callback that receives a BatchProgress record after each completed item. | None |
output_path | Optional[Union[str, Path]] | Optional atomically written final result file. | None |
checkpoint_path | Optional[Union[str, Path]] | Optional PHI-free durable checkpoint file. | None |
resume_from_checkpoint | bool | Resume the committed result prefix. | False |
checkpoint_interval | Optional[int] | Per-run override for checkpoint frequency. | None |
output_format | str |
| 'json' |
Returns:
| Type | Description |
|---|---|
BatchResult | BatchResult with all processing results. |
process_files(file_paths, encoding='utf-8', progress_callback=None, *, on_progress=None, output_path=None, checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, output_format='json') ¶
Process multiple files.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
file_paths | Sequence[Union[str, Path]] | Paths to text files. | required |
encoding | str | File encoding. | 'utf-8' |
progress_callback | Optional[ProgressCallback] | Optional callback for progress updates. | None |
on_progress | Optional[BatchProgressCallback] | Optional PHI-safe callback that receives a BatchProgress record after each completed item. | None |
output_path | Optional[Union[str, Path]] | Optional atomically written final result file. | None |
checkpoint_path | Optional[Union[str, Path]] | Optional PHI-free durable checkpoint file. | None |
resume_from_checkpoint | bool | Resume the committed result prefix. | False |
checkpoint_interval | Optional[int] | Per-run override for checkpoint frequency. | None |
output_format | str |
| 'json' |
Returns:
| Type | Description |
|---|---|
BatchResult | BatchResult with all processing results. |
process_files_to_directory(file_paths, *, input_root, output_dir, encoding='utf-8', checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, progress_callback=None, on_progress=None) ¶
De-identify files into an atomic, checkpointed output directory.
Output paths preserve each input's location relative to input_root. Committed files are hashed in the PHI-free checkpoint and verified before a resumed run skips them.
process_items(items, progress_callback=None, *, on_progress=None, output_path=None, checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, output_format='json') ¶
Process a sequence of BatchItem objects.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
items | Sequence[BatchItem] | Sequence of BatchItem objects. | required |
progress_callback | Optional[ProgressCallback] | Optional callback for progress updates. | None |
on_progress | Optional[BatchProgressCallback] | Optional PHI-safe callback that receives a BatchProgress record after each completed item. | None |
output_path | Optional[Union[str, Path]] | Optional atomically written final result file. | None |
checkpoint_path | Optional[Union[str, Path]] | Optional PHI-free durable checkpoint file. | None |
resume_from_checkpoint | bool | Resume the committed result prefix. | False |
checkpoint_interval | Optional[int] | Per-run override for checkpoint frequency. | None |
output_format | str |
| 'json' |
Returns:
| Type | Description |
|---|---|
BatchResult | BatchResult with all processing results. |
process_texts(texts, ids=None, progress_callback=None, *, on_progress=None, output_path=None, checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, output_format='json') ¶
Process multiple texts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
texts | Sequence[str] | Sequence of texts to analyze. | required |
ids | Optional[Sequence[str]] | Optional identifiers for each text. | None |
progress_callback | Optional[ProgressCallback] | Optional callback for progress updates. Signature: callback(completed_count, total_count, result) | None |
on_progress | Optional[BatchProgressCallback] | Optional PHI-safe callback that receives a BatchProgress record after each completed item. | None |
output_path | Optional[Union[str, Path]] | Optional atomically written final result file. | None |
checkpoint_path | Optional[Union[str, Path]] | Optional PHI-free durable checkpoint file. | None |
resume_from_checkpoint | bool | Resume the committed result prefix instead of starting a new checkpoint. | False |
checkpoint_interval | Optional[int] | Per-run override for checkpoint frequency. | None |
output_format | str |
| 'json' |
Returns:
| Type | Description |
|---|---|
BatchResult | BatchResult with all processing results. |
resume_from_checkpoint(items, *, checkpoint_path, output_path=None, progress_callback=None, on_progress=None, output_format='json') ¶
Resume items from a previously committed batch checkpoint.
PIIEntity¶
Bases: EntityPrediction
Extended Entity with PII-specific metadata.
Attributes:
| Name | Type | Description |
|---|---|---|
text | str | The entity text span |
label | str | PII category (NAME, EMAIL, PHONE, etc.) |
start | Optional[int] | Character start position |
end | Optional[int] | Character end position |
confidence | float | Model confidence score (0-1) |
entity_type | str | PII category (same as label) |
redacted_text | Optional[str] | Replacement text after de-identification |
original_text | Optional[str] | Original text before redaction |
hash_value | Optional[str] | Consistent hash for entity linking |
reversible_id | Optional[str] | Optional reversible pseudonymization handle |
__post_init__() ¶
Initialize entity_type from label if not set.
DeidentificationResult¶
Result of de-identification operation.
Attributes:
| Name | Type | Description |
|---|---|---|
original_text | str | Input text before de-identification |
deidentified_text | str | Output text with PII redacted |
pii_entities | list[PIIEntity] | List of detected and redacted PII entities |
method | str | De-identification method used |
timestamp | datetime | When de-identification was performed |
mapping | Optional[dict[str, str]] | Optional mapping for re-identification. Colliding replacement surfaces use private occurrence keys so separate source spellings remain reversible without changing the de-identified text. |
to_dataframe() ¶
Convert detected PII entities to a pandas DataFrame.
Returns:
| Type | Description |
|---|---|
Any | A pandas DataFrame with one row per detected entity and columns |
Any |
|
Any |
|
Raises:
| Type | Description |
|---|---|
MissingExtraError | If pandas is not installed. This remains an :class: |
to_dict() ¶
Convert result to dictionary format.
Returns:
| Type | Description |
|---|---|
dict | Dictionary with all result fields and metadata |
PDF redaction and fidelity¶
Render and verify a clean redacted PDF from projected rectangles.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source | str | Path | Source digital PDF. Processing is local and never uses a network. | required |
output | str | Path | Destination PDF. It is published atomically only after every mandatory verification passes. | required |
regions | Iterable[Any] | Iterable of | required |
render_dpi | int | Resolution used to burn each source page into safe pixels. | _DEFAULT_RENDER_DPI |
fidelity_dpi | int | None | Resolution used by the independent regression check. Defaults to | None |
pixel_tolerance | int | Maximum per-channel 0-255 difference treated as stable. | _DEFAULT_PIXEL_TOLERANCE |
max_outside_changed_fraction | float | Maximum fraction of pixels outside all redaction boxes that may differ. | _DEFAULT_MAX_OUTSIDE_CHANGED_FRACTION |
max_pages | int | Maximum number of source pages accepted for one render. | _DEFAULT_MAX_PAGES |
max_page_pixels | int | Maximum rendered pixels accepted for any page. | _DEFAULT_MAX_PAGE_PIXELS |
max_total_pixels | int | Maximum rendered pixels accepted across all pages. | _DEFAULT_MAX_TOTAL_PIXELS |
max_regions | int | Maximum number of distinct redaction rectangles. | _DEFAULT_MAX_REGIONS |
overwrite | bool | Permit atomically replacing an existing output file. | False |
Returns:
| Name | Type | Description |
|---|---|---|
A | PdfRedactionResult | class: |
Raises:
| Type | Description |
|---|---|
ValueError | If regions, thresholds, page indexes, or bboxes are invalid. |
FileExistsError | If |
PdfRenderVerificationError | If text removal, boxes, or layout fail. |
MissingDependencyError | If the |
Verify that redacted scrubbed the PHI spans present in original.
spans describes what was redacted. Each item may be:
- a
(page, bbox)region (a mapping/object withpageandbbox, or a :class:~openmed.multimodal.documents_pdf.ProjectedRectangle), or - a character span into
original's extracted text (a(start, end)tuple, or a mapping/object withstart/end) that is projected to a page rectangle usingoriginal.
For each region the verifier asserts (a) no selectable word in redacted remains under the region and (b) an opaque redaction box covers the region. When a raster backend is available (or rasterizer is supplied) the region is also rendered to pixels in both documents and required to differ.
Returns a :class:PdfFidelityReport. Pass strict=True to raise :class:RedactionFidelityError on any residual leakage instead.
Verify that every selected source-word occurrence was removed.
Unlike :func:verify_redacted_pdf, which checks the output text layer at each projected rectangle, this helper accounts for selected word occurrences across the complete extracted text of both PDFs. It therefore catches source words that were moved, reordered, separated, split, merged, or duplicated during re-rendering while allowing unrelated identical words to remain. Reports contain only geometry, counts, and SHA-256 digests; source and residual plaintext are never stored.
A region with no extractable source text fails closed because this helper cannot prove removal. Scanned/image-only PDFs require the separate OCR path. Pass strict=True to raise :class:RedactedTextRemovalError on failure.
Measure page geometry and pixels outside redaction rectangles.
The comparison is deterministic for a fixed local PDF stack. Pixels inside each requested rectangle, plus a one-point antialiasing margin, are masked out. The remaining pixels form an enforceable regression gate. No OCR or page text is included in the returned report.
PipelineTelemetry¶
Opt-in OpenTelemetry spans and metrics for core pipeline stages.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
enabled | bool | Explicit opt-in. The default is | False |
tracer | Any | Optional caller-owned OpenTelemetry tracer. When omitted after opt-in, the global API tracer is used if OpenTelemetry is installed. | None |
meter | Any | Optional caller-owned OpenTelemetry meter. When omitted after opt-in, the global API meter is used if OpenTelemetry is installed. | None |
OpenMed never configures the providers behind these objects and therefore never creates an exporter or network destination.
StageTelemetry¶
No-PHI recorder for one pipeline stage.
Instances are created by :meth:PipelineTelemetry.stage_span. Setters accept only aggregate values and route every span attribute through :func:safe_stage_attributes.
active property ¶
Return whether this recorder has a trace or metric sink.
finish(duration_ms) ¶
Finish the stage and record its duration and aggregate metrics.
mark_failed() ¶
Mark a failed stage without recording an exception or message.
set_entity_count(count) ¶
Record how many entities this stage produced.
set_input_length(length) ¶
Record a stage input character count.
set_labels(labels) ¶
Record a set of canonical category labels, never entity text.
set_offset_range(start, end) ¶
Record aggregate output bounds without storing a detected surface.
set_redacted_length(length) ¶
Record the emitted redacted character count.
set_span_count(count) ¶
Record how many canonical spans this stage produced.