Skip to content

API Reference

extract_pii

Extract PII entities from text with intelligent entity merging.

Uses token classification models to detect personally identifiable information including names, emails, phone numbers, addresses, and other HIPAA-protected identifiers.

The smart merging feature uses regex patterns to identify semantic units (dates, SSN, phone numbers, etc.) and merges fragmented model predictions into complete entities with dominant label selection.

Parameters:

Name Type Description Default
text str | bytes | bytearray | memoryview

Input text to analyze

required
model_name str

PII detection model (registry key or HuggingFace ID). When the default is used and lang is not "en", the language-appropriate default model is selected automatically.

_DEFAULT_EN_MODEL
confidence_threshold float

Minimum confidence score (0-1)

0.5
config Optional[OpenMedConfig]

Optional configuration override

None
use_smart_merging bool

Enable regex-based semantic unit merging (recommended)

True
lang str

ISO 639-1 language code (en, fr, de, it, es, nl, hi, te, pt, ar, ja, tr). Controls which default model and regex patterns are used. Mixed Latin/Devanagari or Latin/Telugu notes automatically use script-aware India clinical routing for hi and te.

'en'
normalize_accents Optional[bool]

Strip diacritical marks before model inference so that models trained on accent-free text still detect accented names. Entity spans in the result reference the original (accented) text. None (default) auto-enables for languages in _ACCENT_NORMALIZE_LANGS (currently Spanish).

None
preserve_whitespace bool

Preserve leading and trailing source whitespace so returned entity offsets refer to the exact input string.

False
loader Optional['ModelLoader']

Optional shared model loader to reuse warmed pipelines.

None
batch_size Optional[int]

Optional backend inference batch size.

None
num_workers Optional[int]

Optional backend inference worker count.

None
custom_recognizer Any

Optional deny-list/allow-list recognizer config, CustomRecognizer instance, or JSON/YAML config path. Deny-list matches are added with custom:deny provenance; allow-list matches suppress overlapping spans from any detector.

None
abdm Optional[bool]

Enable the India ABDM identifier bundle. None auto-enables it for Hindi/Telugu and India locales; False explicitly disables that automatic activation.

None
code_mixed bool

Enable the explicit English/Hinglish route. This preserves English model detection while adding Roman-Hindi context patterns.

False
token_language_tags Optional[Sequence[Any]]

Required with code_mixed=True. Ordered, non-overlapping records with start, end, and label (en, hi, ne, univ, or other). Tags are consumed as offsets/labels only and are never copied with raw token surfaces. When omitted in code-mixed mode, the deterministic token-LID fallback derives them from the input.

None
lid_model Optional['TokenLIDHook']

Optional user-supplied token language-ID hook used only when code_mixed=True and token_language_tags is omitted.

None
transliterated_name_config Any

Optional configuration for the conservative Latin-script Indian given/family-name allow/deny bridge.

None
cache_results bool

Whether to cache this result in the in-process LRU cache. Cached results may contain PHI, but are never saved to disk.

False
max_cache_entries int

Maximum number of cached results.

128
budget Optional[RequestBudget]

Optional per-request wall-time and input-character budget. None means unlimited. Deadline checks are cooperative, and an over-length input is rejected before model inference.

None

Returns:

Type Description
PredictionResult

PredictionResult with detected PII entities

Raises:

Type Description
InputError

If text or request options are malformed or unsupported.

CapabilityError

If the selected model or optional runtime is unavailable.

BudgetExceededError

If the request exceeds its configured budget.

InternalError

If inference violates a result or span invariant.

Example

from unittest.mock import patch from openmed.core.pii import extract_pii from openmed.processing.outputs import EntityPrediction, PredictionResult fake_result = PredictionResult( ... text="Patient Casey Example called.", ... entities=[ ... EntityPrediction( ... text="Casey Example", ... label="NAME", ... confidence=0.98, ... start=8, ... end=21, ... ) ... ], ... model_name="fixture-pii-model", ... timestamp="2026-01-01T00:00:00", ... ) with patch("openmed.analyze_text", return_value=fake_result): ... result = extract_pii( ... "Patient Casey Example called.", ... model_name="fixture-pii-model", ... use_smart_merging=False, ... ) next((entity.text, entity.label) for entity in result.entities) ('Casey Example', 'NAME')

deidentify

De-identify text by detecting and redacting PII with intelligent merging.

Implements multiple de-identification strategies for HIPAA compliance:

  • mask: Replace with placeholders like [NAME], [EMAIL], etc.
  • aadhaar_mask: Render valid Aadhaar values as XXXX XXXX NNNN; use ordinary placeholders for every other entity
  • remove: Remove PII text entirely (empty string)
  • replace: Replace with fake but realistic data
  • hash: Replace with consistent hashed values for entity linking
  • format_preserve: Replace structured identifiers with synthetic values that keep shape and separators, masking unsupported labels
  • shift_dates: Shift dates by random offset while preserving intervals

Smart merging uses regex patterns to merge fragmented entities (e.g., dates split into '01' and '/15/1970' are merged into complete '01/15/1970').

Code-mixed mode is explicit and offset driven. With code_mixed=True and per-token language tags, the English NER path remains active while a separate Roman-script Hindi pattern bank detects cues such as naam, umar, pata, mobile, and janm. The combined spans pass through the normal entity merger and final safety sweep before redaction.

Parameters:

Name Type Description Default
text str | bytes | bytearray | memoryview

Input text to de-identify

required
method DeidentificationMethod

De-identification method (mask, aadhaar_mask, remove, replace, hash, shift_dates, format_preserve)

'mask'
model_name str

PII detection model

_DEFAULT_EN_MODEL
confidence_threshold float

Minimum confidence for redaction (default 0.7 for safety)

0.7
keep_year bool

For dates, keep the year unchanged

False
shift_dates Optional[bool]

Deprecated alias for method="shift_dates".

None
date_shift_days Optional[int]

Specific number of days to shift when patient_key is omitted. When patient_key is supplied, this is treated as a legacy maximum absolute offset bound unless date_shift_max_days is also supplied.

None
patient_key Optional[str | bytes]

Optional stable patient identifier used only to derive a deterministic HMAC date-shift offset. Raw keys are not logged, persisted, or returned.

None
date_shift_max_days Optional[int]

Maximum absolute offset for random, seeded, or patient-keyed date shifting. Defaults to 365 when patient_key or seed is supplied and neither this nor date_shift_days is set.

None
date_shift_secret Optional[str | bytes]

Required HMAC key material for patient-keyed offsets. Reuse the same value across sessions to keep offsets stable.

None
keep_mapping bool

Keep mapping for re-identification

False
config Optional[OpenMedConfig]

Optional configuration override

None
use_smart_merging bool

Enable regex-based semantic unit merging (recommended)

True
use_safety_sweep bool

Run a deterministic structured-identifier sweep after model detection and before redaction.

True
lang str

ISO 639-1 language code (en, fr, de, it, es, nl, hi, te, pt, ar, ja, tr). Controls model selection, regex patterns, and fake data for replacement. Mixed Latin/Devanagari or Latin/Telugu notes automatically use the script-aware India clinical route for hi and te.

'en'
normalize_accents Optional[bool]

Strip diacritical marks before model inference. None (default) auto-enables for Spanish.

None
loader Optional['ModelLoader']

Optional shared model loader to reuse warmed pipelines.

None
consistent bool

When method="replace" or method="format_preserve", generate stable surrogates (same input -> same surrogate within the call). Lets repeated mentions of the same name resolve to one fake identity instead of a different one each time.

False
seed Optional[int]

Optional request-scoped integer seed for cross-run reproducibility of replacements and automatic date shifting. Implies consistent=True for replacement methods. Explicit date_shift_days and patient-keyed offsets still take precedence for method="shift_dates".

None
locale Optional[str]

Faker locale override (pt_BR, en_GB, ...) for method="replace" and method="format_preserve". When None, derived from lang.

None
surrogate_vault Optional['SurrogateVault']

Optional cross-document surrogate vault. When provided with method="replace", OpenMed stores only HMAC source hashes. Indian names in Devanagari, Tamil, or opted-in Romanization use one HMAC of an in-memory phonetic fold and render the reused identity in the input script; the fold itself is never persisted or audited.

None
policy Optional[str]

Optional policy profile name controlling arbitration, action selection, mandatory safety sweep behavior, and reversible mapping.

None
calibration_thresholds_path Optional[str | Path]

Optional thresholds.json artifact path or artifact directory. When provided, per-label calibrated thresholds filter model detections and appear in audit output.

None
custom_recognizer Any

Optional deny-list/allow-list recognizer config, CustomRecognizer instance, or JSON/YAML config path. Deny-list matches are redacted with custom:deny provenance; allow-list matches suppress overlapping spans from any detector.

None
abdm Optional[bool]

Enable the India ABDM recognizer bundle. None auto-enables it for policy="india_dpdp_act", Hindi/Telugu, or an India locale. Pass False to opt out of automatic activation.

None
code_mixed bool

Enable the explicit English/Hinglish de-identification path.

False
token_language_tags Optional[Sequence[Any]]

Required with code_mixed=True. Ordered token records with exact start/end offsets and an en, hi, ne, univ, or other label. A pure-English tag stream does not activate Roman-Hindi patterns. When omitted in code-mixed mode, the deterministic token-LID fallback derives the tags.

None
lid_model Optional['TokenLIDHook']

Optional user-supplied token language-ID hook used only when code-mixed tags are inferred.

None
transliterated_name_config Any

Optional configuration for the Latin-script Indian name allow/deny bridge. The default bridge is conservative and can be replaced or extended by configuration.

None
audit bool

Return an AuditReport instead of the DeidentificationResult. Fresh calls use separate private HMAC keys. For stable hashes across runs, use a Pipeline with an explicit private hmac_secret.

False
cache_results bool

Whether to cache this result in the in-process LRU cache. Cached results may contain PHI, but are never saved to disk.

False
max_cache_entries int

Maximum number of cached results.

128
budget Optional[RequestBudget]

Optional per-request wall-time and input-character budget. None means unlimited. Deadline checks are cooperative, and an over-length input is rejected before model inference.

None

Returns:

Type Description
DeidentificationResult | 'AuditReport'

DeidentificationResult with original and de-identified text, or

DeidentificationResult | 'AuditReport'

AuditReport when audit=True.

Raises:

Type Description
InputError

If text or de-identification options are malformed.

PolicyError

If the selected privacy policy is invalid or disallows the requested operation.

CapabilityError

If the selected model or optional runtime is unavailable.

BudgetExceededError

If the request exceeds its configured budget.

InternalError

If inference or redaction violates an internal invariant.

Example

from datetime import datetime from types import SimpleNamespace from unittest.mock import patch from openmed.core.pii import ( ... DeidentificationResult, ... PIIEntity, ... deidentify, ... ) fixture = DeidentificationResult( ... original_text="Patient Casey Example", ... deidentified_text="Patient [NAME]", ... pii_entities=[ ... PIIEntity( ... text="Casey Example", ... label="NAME", ... start=8, ... end=21, ... confidence=0.98, ... redacted_text="[NAME]", ... ) ... ], ... method="mask", ... timestamp=datetime(2026, 1, 1, 0, 0, 0), ... mapping={"[NAME]": "Casey Example"}, ... ) with patch("openmed.core.pipeline.Pipeline") as pipeline_cls: ... pipeline_cls.return_value.run.return_value = SimpleNamespace( ... deidentification_result=fixture ... ) ... result = deidentify( ... "Patient Casey Example", ... method="mask", ... keep_mapping=True, ... ) result.deidentified_text 'Patient [NAME]' result.mapping

reidentify

Re-identify text using stored mapping.

Restores original PII from de-identified text using the mapping created during de-identification. Only works if keep_mapping=True was used.

Parameters:

Name Type Description Default
deidentified_text str

De-identified text

required
mapping Mapping[str, str]

Mapping from redacted to original text, optionally including occurrence-aware entries emitted for colliding replacement values.

required

Returns:

Type Description
str

Re-identified text with original PII restored

Raises:

Type Description
InputError

If the text or mapping does not use the documented string types.

Example

from openmed.core.pii import reidentify reidentify( ... "Patient [NAME] has record [ID]", ... {"[NAME]": "Casey Example", "[ID]": "MRN-0001"}, ... ) 'Patient Casey Example has record MRN-0001'

Note

Only works if keep_mapping=True was used during de-identification. Requires proper authorization and audit logging in production.

Structured errors

The complete hierarchy, compatibility guarantees, and REST/MCP mappings are documented in Structured public errors.

OpenMedError

Bases: Exception

Base class for expected failures on the OpenMed public API.

Parameters:

Name Type Description Default
message str

Actionable, PHI-free explanation of the failure.

required
code Optional[str]

Optional stable leaf code owned by OpenMed. Callers should not invent codes; this hook supports existing specialized errors.

None
details Optional[Mapping[str, Any]]

Optional PHI-free structured context.

None

Attributes:

Name Type Description
code str

Stable machine-readable error code.

message

Actionable, PHI-free human-readable message.

details dict[str, Any]

PHI-free structured context.

__str__()

Return the PHI-free human-readable message.

to_dict(*, include_details=True)

Return a JSON-ready error object.

Parameters:

Name Type Description Default
include_details bool

Include the structured details mapping. Service adapters use False for server-side failures.

True

Returns:

Type Description
dict[str, Any]

A mapping with stable code and actionable message fields.

InputError

Bases: OpenMedError, ValueError, TypeError

Caller input is malformed, conflicting, or unsupported.

ValueError and TypeError remain bases so existing handlers keep catching value and type validation failures after adopting the taxonomy.

ConfigurationError

Bases: OpenMedError, ValueError, TypeError, KeyError

Configuration is missing, unknown, or inconsistent.

The legacy ValueError, TypeError, and KeyError bases preserve compatibility with configuration registries and validators.

CapabilityError

Bases: OpenMedError, ImportError

A requested runtime, model, or optional capability is unavailable.

MissingExtraError

Bases: CapabilityError

An optional package or OpenMed extra is not installed.

Parameters:

Name Type Description Default
message str

Actionable message containing an installation instruction.

required
package Optional[str]

Missing distribution name, if known.

None
feature Optional[str]

Feature that requires the package, if known.

None
extra Optional[str]

OpenMed extra that provides the package, if known.

None
details Optional[Mapping[str, Any]]

Additional PHI-free structured context.

None

ModelLoadError

Bases: CapabilityError, ValueError

A model, tokenizer, or inference backend could not be loaded.

ValueError is retained in addition to ImportError because older model-loading paths used both builtin families.

Parameters:

Name Type Description Default
message str

Actionable, PHI-free load failure message.

required
model_name Optional[str]

Non-sensitive model identifier or local path, if known.

None
details Optional[Mapping[str, Any]]

Additional PHI-free structured context.

None

PolicyError

Bases: OpenMedError, ValueError, TypeError

A request violates or misconfigures a privacy policy constraint.

BudgetExceededError

Bases: OpenMedError, RuntimeError

A request exceeded a configured size, time, or resource budget.

This class accepts both the generic taxonomy constructor and the historical request-budget fields used by :mod:openmed.core.budget.

Parameters:

Name Type Description Default
message Optional[str]

Optional actionable message. When omitted, one is built from kind, limit, observed, and checkpoint.

None
kind Optional[str]

Budget dimension such as "wall_time" or "input_chars".

None
limit Optional[float]

Configured limit.

None
observed Optional[float]

Observed value that exceeded the limit.

None
checkpoint Optional[str]

Safe pipeline checkpoint where the limit was observed.

None
details Optional[Mapping[str, Any]]

Additional PHI-free structured context.

None

InternalError

Bases: OpenMedError, RuntimeError

An internal invariant failed and the request cannot safely continue.

InferenceError

Bases: InternalError

A model or backend returned a structurally invalid inference result.

redact_detail

Return a stable descriptor for untrusted text without exposing it.

Parameters:

Name Type Description Default
value Any

Value to describe. It is converted to text only for hashing and is never included verbatim in the returned descriptor.

required

Returns:

Type Description
str

A descriptor containing only the UTF-8 byte length and SHA-256 digest.

analyze_text

Run a token-classification model on text and format the predictions.

Parameters:

Name Type Description Default
text str

Clinical or biomedical text to analyse.

required
model_name str

Registry key, fully-qualified Hugging Face model id, or local model path.

'disease_detection_superclinical'
model_id Optional[str]

Alias for model_name. Useful for APIs and examples that name model identifiers as model_id.

None
config Optional[OpenMedConfig]

Optional :class:~openmed.core.config.OpenMedConfig instance.

None
loader Optional[ModelLoader]

Reuse an existing :class:~openmed.core.models.ModelLoader.

None
aggregation_strategy Optional[str]

Hugging Face aggregation strategy ("simple" by default). Set to None to work with raw token outputs.

'simple'
output_format str

"dict" (default), "json", "html" or "csv".

'dict'
include_confidence bool

Whether to include confidence scores in formatted output.

True
confidence_threshold Optional[float]

Minimum confidence for entities. None keeps all.

0.0
group_entities bool

Merge adjacent entities of the same label in the formatted output.

False
formatter_kwargs Optional[Dict[str, Any]]

Extra keyword arguments forwarded to :func:openmed.processing.format_predictions.

None
metadata Optional[Dict[str, Any]]

Optional metadata to attach to the result.

None
use_fast_tokenizer bool

Prefer fast tokenizers when available.

True
sentence_detection bool

Enable sentence detection (default: True). The engine is selected by sentence_backend.

True
sentence_language str

Language hint for the sentence detector.

'en'
sentence_clean bool

Whether to enable the sentence detector's cleaning heuristics.

False
sentence_segmenter Optional[Any]

Optional preconstructed segmenter object to reuse. It cannot be combined with sentence_backend="yasbd".

None
sentence_backend Literal['auto', 'yasbd']

Sentence segmentation engine to use. It can be "auto" (default, unchanged routing) or "yasbd" (experimental opt-in; requires openmed[yasbd]).

'auto'
assert_context bool

Attach deterministic negation, uncertainty, experiencer, and temporality labels to each entity under metadata["clinical_context"]. Disabled by default.

False
cache_results bool

Whether to cache this result in the in-process LRU cache. Cached results may contain PHI, but are never saved to disk.

False
max_cache_entries int

Maximum number of cached results.

128
**pipeline_kwargs Any

Additional arguments passed to :meth:openmed.core.models.ModelLoader.create_pipeline.

{}

Returns:

Type Description
Union[AnalyzeResult, str, List[Dict[str, Any]]]

Analyze result for "dict" output, otherwise the requested rendered

Union[AnalyzeResult, str, List[Dict[str, Any]]]

format.

Example

class FixtureLoader: ... config = None ... ... def create_pipeline(self, model_name, kwargs): ... def pipeline(text, call_kwargs): ... return [ ... { ... "entity_group": "CONDITION", ... "score": 0.99, ... "start": 11, ... "end": 17, ... "word": "asthma", ... } ... ] ... ... return pipeline ... ... def get_max_sequence_length(self, model_name, tokenizer=None): ... return 128 result = analyze_text( ... "History of asthma.", ... model_name="fixture-ner-model", ... loader=FixtureLoader(), ... sentence_detection=False, ... ) next((entity.text, entity.label) for entity in result.entities) ('asthma', 'CONDITION')

list_models

Return available OpenMed model identifiers.

Parameters:

Name Type Description Default
include_registry bool

Include entries from the bundled registry in addition to entries in the committed manifest.

True
include_remote bool

Retained for compatibility; no live discovery is performed.

True
config Optional[OpenMedConfig]

Optional custom configuration for model discovery.

None

BatchProcessor

Process multiple texts efficiently with progress tracking.

Example usage

from openmed import BatchProcessor, OpenMedConfig processor = BatchProcessor(model_name="disease_detection_superclinical") texts = ["Patient has diabetes.", "No significant findings."] result = processor.process_texts(texts) print(result.summary())

__init__(model_name='disease_detection_superclinical', operation='analyze_text', batch_size=8, config=None, loader=None, aggregation_strategy='simple', confidence_threshold=None, group_entities=False, continue_on_error=True, checkpoint_interval=_DEFAULT_CHECKPOINT_INTERVAL, _atomic_write_hook=None, budget=None, **analyze_kwargs)

Initialize batch processor.

Parameters:

Name Type Description Default
model_name str

Model registry key or HuggingFace identifier.

'disease_detection_superclinical'
operation BatchOperation

Which function to call per item: "analyze_text" (default), "extract_pii" or "deidentify". Extra kwargs passed via **analyze_kwargs are passed to the selected function.

'analyze_text'
batch_size int

Number of documents to process together per batch.

8
config Optional[Any]

Optional OpenMedConfig instance.

None
loader Optional[Any]

Optional ModelLoader instance to reuse.

None
aggregation_strategy Optional[str]

HuggingFace aggregation strategy (analyze_text operation only).

'simple'
confidence_threshold Optional[float]

Minimum confidence for entities. When not provided, defaults match the selected operation: 0.0 for analyze_text, 0.5 for extract_pii, and 0.7 for deidentify.

None
group_entities bool

Whether to group adjacent entities (analyze_text operation only).

False
continue_on_error bool

Continue processing on individual item errors.

True
checkpoint_interval int

Maximum number of items processed between durable checkpoints.

_DEFAULT_CHECKPOINT_INTERVAL
budget Optional[Any]

Optional per-request resource budget applied independently to each extract_pii or deidentify item. None means unlimited. The analyze_text operation ignores this value.

None
**analyze_kwargs Any

Additional arguments passed to the selected function.

{}

iter_process(texts, ids=None, *, on_progress=None)

Process texts as an iterator, yielding results one at a time.

This is useful for streaming results or processing very large batches where you don't want to hold all results in memory.

Parameters:

Name Type Description Default
texts Sequence[str]

Sequence of texts to analyze.

required
ids Optional[Sequence[str]]

Optional identifiers for each text.

None
on_progress Optional[BatchProgressCallback]

Optional PHI-safe callback that receives a BatchProgress record after each completed item.

None

Yields:

Type Description
BatchItemResult

BatchItemResult for each processed text.

process_directory(directory, pattern='*.txt', recursive=False, encoding='utf-8', progress_callback=None, *, on_progress=None, output_path=None, checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, output_format='json')

Process all matching files in a directory.

Parameters:

Name Type Description Default
directory Union[str, Path]

Directory path.

required
pattern str

Glob pattern for file matching.

'*.txt'
recursive bool

Whether to search recursively.

False
encoding str

File encoding.

'utf-8'
progress_callback Optional[ProgressCallback]

Optional callback for progress updates.

None
on_progress Optional[BatchProgressCallback]

Optional PHI-safe callback that receives a BatchProgress record after each completed item.

None
output_path Optional[Union[str, Path]]

Optional atomically written final result file.

None
checkpoint_path Optional[Union[str, Path]]

Optional PHI-free durable checkpoint file.

None
resume_from_checkpoint bool

Resume the committed result prefix.

False
checkpoint_interval Optional[int]

Per-run override for checkpoint frequency.

None
output_format str

"json" or "summary" for output_path.

'json'

Returns:

Type Description
BatchResult

BatchResult with all processing results.

process_files(file_paths, encoding='utf-8', progress_callback=None, *, on_progress=None, output_path=None, checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, output_format='json')

Process multiple files.

Parameters:

Name Type Description Default
file_paths Sequence[Union[str, Path]]

Paths to text files.

required
encoding str

File encoding.

'utf-8'
progress_callback Optional[ProgressCallback]

Optional callback for progress updates.

None
on_progress Optional[BatchProgressCallback]

Optional PHI-safe callback that receives a BatchProgress record after each completed item.

None
output_path Optional[Union[str, Path]]

Optional atomically written final result file.

None
checkpoint_path Optional[Union[str, Path]]

Optional PHI-free durable checkpoint file.

None
resume_from_checkpoint bool

Resume the committed result prefix.

False
checkpoint_interval Optional[int]

Per-run override for checkpoint frequency.

None
output_format str

"json" or "summary" for output_path.

'json'

Returns:

Type Description
BatchResult

BatchResult with all processing results.

process_files_to_directory(file_paths, *, input_root, output_dir, encoding='utf-8', checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, progress_callback=None, on_progress=None)

De-identify files into an atomic, checkpointed output directory.

Output paths preserve each input's location relative to input_root. Committed files are hashed in the PHI-free checkpoint and verified before a resumed run skips them.

process_items(items, progress_callback=None, *, on_progress=None, output_path=None, checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, output_format='json')

Process a sequence of BatchItem objects.

Parameters:

Name Type Description Default
items Sequence[BatchItem]

Sequence of BatchItem objects.

required
progress_callback Optional[ProgressCallback]

Optional callback for progress updates.

None
on_progress Optional[BatchProgressCallback]

Optional PHI-safe callback that receives a BatchProgress record after each completed item.

None
output_path Optional[Union[str, Path]]

Optional atomically written final result file.

None
checkpoint_path Optional[Union[str, Path]]

Optional PHI-free durable checkpoint file.

None
resume_from_checkpoint bool

Resume the committed result prefix.

False
checkpoint_interval Optional[int]

Per-run override for checkpoint frequency.

None
output_format str

"json" or "summary" for output_path.

'json'

Returns:

Type Description
BatchResult

BatchResult with all processing results.

process_texts(texts, ids=None, progress_callback=None, *, on_progress=None, output_path=None, checkpoint_path=None, resume_from_checkpoint=False, checkpoint_interval=None, output_format='json')

Process multiple texts.

Parameters:

Name Type Description Default
texts Sequence[str]

Sequence of texts to analyze.

required
ids Optional[Sequence[str]]

Optional identifiers for each text.

None
progress_callback Optional[ProgressCallback]

Optional callback for progress updates. Signature: callback(completed_count, total_count, result)

None
on_progress Optional[BatchProgressCallback]

Optional PHI-safe callback that receives a BatchProgress record after each completed item.

None
output_path Optional[Union[str, Path]]

Optional atomically written final result file.

None
checkpoint_path Optional[Union[str, Path]]

Optional PHI-free durable checkpoint file.

None
resume_from_checkpoint bool

Resume the committed result prefix instead of starting a new checkpoint.

False
checkpoint_interval Optional[int]

Per-run override for checkpoint frequency.

None
output_format str

"json" or "summary" for output_path.

'json'

Returns:

Type Description
BatchResult

BatchResult with all processing results.

resume_from_checkpoint(items, *, checkpoint_path, output_path=None, progress_callback=None, on_progress=None, output_format='json')

Resume items from a previously committed batch checkpoint.

PIIEntity

Bases: EntityPrediction

Extended Entity with PII-specific metadata.

Attributes:

Name Type Description
text str

The entity text span

label str

PII category (NAME, EMAIL, PHONE, etc.)

start Optional[int]

Character start position

end Optional[int]

Character end position

confidence float

Model confidence score (0-1)

entity_type str

PII category (same as label)

redacted_text Optional[str]

Replacement text after de-identification

original_text Optional[str]

Original text before redaction

hash_value Optional[str]

Consistent hash for entity linking

reversible_id Optional[str]

Optional reversible pseudonymization handle

__post_init__()

Initialize entity_type from label if not set.

DeidentificationResult

Result of de-identification operation.

Attributes:

Name Type Description
original_text str

Input text before de-identification

deidentified_text str

Output text with PII redacted

pii_entities list[PIIEntity]

List of detected and redacted PII entities

method str

De-identification method used

timestamp datetime

When de-identification was performed

mapping Optional[dict[str, str]]

Optional mapping for re-identification. Colliding replacement surfaces use private occurrence keys so separate source spellings remain reversible without changing the de-identified text.

to_dataframe()

Convert detected PII entities to a pandas DataFrame.

Returns:

Type Description
Any

A pandas DataFrame with one row per detected entity and columns

Any

text, label, entity_type, start, end,

Any

confidence, action, and result_id.

Raises:

Type Description
MissingExtraError

If pandas is not installed. This remains an :class:ImportError for compatibility.

to_dict()

Convert result to dictionary format.

Returns:

Type Description
dict

Dictionary with all result fields and metadata

PDF redaction and fidelity

Render and verify a clean redacted PDF from projected rectangles.

Parameters:

Name Type Description Default
source str | Path

Source digital PDF. Processing is local and never uses a network.

required
output str | Path

Destination PDF. It is published atomically only after every mandatory verification passes.

required
regions Iterable[Any]

Iterable of (page, bbox) tuples, mappings/objects with page and bbox, or ProjectedRectangle instances. Bboxes use pdfplumber's top-origin (x0, top, x1, bottom) coordinates.

required
render_dpi int

Resolution used to burn each source page into safe pixels.

_DEFAULT_RENDER_DPI
fidelity_dpi int | None

Resolution used by the independent regression check. Defaults to render_dpi.

None
pixel_tolerance int

Maximum per-channel 0-255 difference treated as stable.

_DEFAULT_PIXEL_TOLERANCE
max_outside_changed_fraction float

Maximum fraction of pixels outside all redaction boxes that may differ.

_DEFAULT_MAX_OUTSIDE_CHANGED_FRACTION
max_pages int

Maximum number of source pages accepted for one render.

_DEFAULT_MAX_PAGES
max_page_pixels int

Maximum rendered pixels accepted for any page.

_DEFAULT_MAX_PAGE_PIXELS
max_total_pixels int

Maximum rendered pixels accepted across all pages.

_DEFAULT_MAX_TOTAL_PIXELS
max_regions int

Maximum number of distinct redaction rectangles.

_DEFAULT_MAX_REGIONS
overwrite bool

Permit atomically replacing an existing output file.

False

Returns:

Name Type Description
A PdfRedactionResult

class:PdfRedactionResult containing PHI-safe verification evidence.

Raises:

Type Description
ValueError

If regions, thresholds, page indexes, or bboxes are invalid.

FileExistsError

If output exists and overwrite is false.

PdfRenderVerificationError

If text removal, boxes, or layout fail.

MissingDependencyError

If the multimodal PDF stack is unavailable.

Verify that redacted scrubbed the PHI spans present in original.

spans describes what was redacted. Each item may be:

  • a (page, bbox) region (a mapping/object with page and bbox, or a :class:~openmed.multimodal.documents_pdf.ProjectedRectangle), or
  • a character span into original's extracted text (a (start, end) tuple, or a mapping/object with start/end) that is projected to a page rectangle using original.

For each region the verifier asserts (a) no selectable word in redacted remains under the region and (b) an opaque redaction box covers the region. When a raster backend is available (or rasterizer is supplied) the region is also rendered to pixels in both documents and required to differ.

Returns a :class:PdfFidelityReport. Pass strict=True to raise :class:RedactionFidelityError on any residual leakage instead.

Verify that every selected source-word occurrence was removed.

Unlike :func:verify_redacted_pdf, which checks the output text layer at each projected rectangle, this helper accounts for selected word occurrences across the complete extracted text of both PDFs. It therefore catches source words that were moved, reordered, separated, split, merged, or duplicated during re-rendering while allowing unrelated identical words to remain. Reports contain only geometry, counts, and SHA-256 digests; source and residual plaintext are never stored.

A region with no extractable source text fails closed because this helper cannot prove removal. Scanned/image-only PDFs require the separate OCR path. Pass strict=True to raise :class:RedactedTextRemovalError on failure.

Measure page geometry and pixels outside redaction rectangles.

The comparison is deterministic for a fixed local PDF stack. Pixels inside each requested rectangle, plus a one-point antialiasing margin, are masked out. The remaining pixels form an enforceable regression gate. No OCR or page text is included in the returned report.

PipelineTelemetry

Opt-in OpenTelemetry spans and metrics for core pipeline stages.

Parameters:

Name Type Description Default
enabled bool

Explicit opt-in. The default is False.

False
tracer Any

Optional caller-owned OpenTelemetry tracer. When omitted after opt-in, the global API tracer is used if OpenTelemetry is installed.

None
meter Any

Optional caller-owned OpenTelemetry meter. When omitted after opt-in, the global API meter is used if OpenTelemetry is installed.

None

OpenMed never configures the providers behind these objects and therefore never creates an exporter or network destination.

disabled() classmethod

Return an explicitly disabled no-op runtime.

from_env(*, tracer=None, meter=None) classmethod

Build telemetry from the explicit environment opt-in.

stage_span(index, name)

Create one fixed-name, no-PHI span for a pipeline stage.

StageTelemetry

No-PHI recorder for one pipeline stage.

Instances are created by :meth:PipelineTelemetry.stage_span. Setters accept only aggregate values and route every span attribute through :func:safe_stage_attributes.

active property

Return whether this recorder has a trace or metric sink.

finish(duration_ms)

Finish the stage and record its duration and aggregate metrics.

mark_failed()

Mark a failed stage without recording an exception or message.

set_entity_count(count)

Record how many entities this stage produced.

set_input_length(length)

Record a stage input character count.

set_labels(labels)

Record a set of canonical category labels, never entity text.

set_offset_range(start, end)

Record aggregate output bounds without storing a detected surface.

set_redacted_length(length)

Record the emitted redacted character count.

set_span_count(count)

Record how many canonical spans this stage produced.