Skip to main content
Glama

Mahrem

Hassas veriyi yapay zekâya gitmeden önce cihazda rumuzlayan açık kaynak gizlilik katmanı.

Mahrem ayrı bir SaaS veya belge yükleme sitesi değildir. Yerel MCP aracı + skill paketi ve bağımsız tarayıcı eklentisi içerir. Kod MIT lisanslıdır, API anahtarı istemez. Claude/GPT kullanımının kendi abonelik veya API maliyeti ayrıdır.

Önemli: Ham metni önce sohbete yapıştırıp sonra “maskele” demek güvenli değildir. Metin o anda modele ulaşmış olur. Mahrem bu nedenle metin değil, yerel dosya yolu kabul eder.

Hukuk profesyonelleri için kolay başlangıç

Windows için indir · Adım adım kurulum rehberi

ZIP dosyasını indirin, Tümünü ayıkla ile açın ve Kur.bat dosyasına çift tıklayın. Gerekli yazılımlar otomatik hazırlanır; Git veya Python komutu yazmanız gerekmez. Kurulum sonunda uygulamanıza gireceğiniz hazır bağlantı bilgileri açılır.

Kurulum çift tıkla başlar; Claude/Codex bağlantısı için rehberdeki kısa ayar adımı da gerekir. Windows kurucusu henüz imzalı değildir. İlk denemeyi paketteki örnek belgeyle yapın.

Tarayıcı kullanıyorsanız: Chrome / Edge eklentisini kurun. Python veya MCP kurulumu gerekmez. Eklenti ayrı bir yerel sekmede metni maskeler; kontrol ettiğiniz çıktıyı sohbete kendiniz kopyalarsınız. Mağazada yayımlanmış değildir; indirilen extension klasörü yüklenir. Sohbet sitelerini izlemez, gönderimi veya dosya yüklemeyi otomatik durdurmaz.

Related MCP server: privacyscrubber-mcp

0.2: Word, PDF ve tarayıcı

  • Yerel araç .txt, .md, Word .docx ve metin içeren .pdf dosyalarını okur.

  • Word/PDF çıktısı maskeli TXT olur. Orijinal dosya değiştirilmez; yeni çıktıda sayfa düzeni, biçim ve dijital imza korunmaz. Bu işlem PDF üzerinde karartma değildir.

  • OCR yoktur. Şifreli PDF ve metin okunamayan sayfa içeren PDF reddedilir. Eski .doc dosyasını önce .docx olarak kaydedin.

  • Görsel, ek ve form içeriği eksik kalabilir. Çıkarılan metni ve maskelenmemiş bilgileri yerelde kontrol edin.

  • Tarayıcı eklentisi yapıştırılan metinle çalışır; doğrudan Word/PDF yüklemez. Yanıtı aynı sekmede geri açıp TXT olarak indirebilirsiniz.

Nasıl çalışır?

flowchart LR
    A[Ham yerel dosya] -->|yalnizca dosya yolu| B[Mahrem MCP]
    B -->|rumuzlu metin| C[Claude / Codex / ChatGPT Desktop]
    C -->|yer tutucular korunur| D[Maskeli sonuc dosyasi]
    D --> B
    B -->|acik metni modele vermeden| E[Geri acilmis yerel dosya]

Ornek:

Musteri: Ayse Deneme
TCKN: 10000000146
E-posta: ayse.deneme@example.com

sununa donusur:

Musteri: [KISI-1]
TCKN: [TCKN-1]
E-posta: [EPOSTA-1]

Aynı değer, aynı maskeleme işlemi içinde aynı rumuzu alır. Her dosya ayrı bir geri açma oturumu oluşturur. Model maskeli metinle çalışır. Sonuç tamamlanınca restore_file, gerçek değerleri yalnızca yerel çıktı dosyasına yazar; açık metni model yanıtına döndürmez.

Tespit kapsamı

  • TCKN: yalnizca checksum'u gecerli adaylar

  • TR IBAN: uzunluk ve mod-97 kontrolu

  • Turkiye cep telefonu

  • E-posta

  • Luhn uyumlu kart numarasi

  • IPv4 ve URL

  • Musteri:, Davaci:, Davali: gibi belirli etiketlerden sonra gelen kisi adlari

  • Yerel rules.json ile ozel kisi/kurum terimleri ve allowlist

  • UTF-8 .txt, .md, .docx ve metin PDF girdileri; TXT/MD çıktısı

Bu bir alpha sürümüdür. Serbest metindeki tüm kişi ve kurum adlarını bulduğunu iddia etmez. Kalıcı şifreli kasa ve tamamen yerel NER modeli henüz yoktur.

Kurulum

Gerekenler: Python 3.10+ ve uv.

git clone https://github.com/orhanluce/mahrem.git
cd mahrem
uv tool install .

mahrem-mcp komutunun PATH uzerinde oldugunu kontrol edin:

mahrem --help

Claude Code

Repo hem Claude plugin manifestini hem de .mcp.json dosyasini icerir:

claude --plugin-dir .

Claude icinde /mahrem:mahrem komutunu kullanabilir veya “bu dosyayi Mahrem ile isle” diyebilirsiniz.

Codex

Yerel MCP sunucusunu bir kez ekleyin:

codex mcp add mahrem -- mahrem-mcp
codex mcp list

Skill'i kullanici kapsaminda kurmak icin skills/mahrem klasorunu %USERPROFILE%\.agents\skills\mahrem altina kopyalayabilirsiniz. Repo bir Codex uyumluluk manifesti de icerir: .codex-plugin/plugin.json.

ChatGPT web yerel Codex yapılandırmasını okumaz. Web sohbeti için tarayıcı eklentisinin ayrı maskeleme akışını kullanın. Windows kurucusu PATH ayarı yapmaz; onunla kurduysanız rehberdeki mutlak komut yolunu kullanın.

Kullanim

Dosya yollarini mutlak verin.

mahrem scan "C:\Belgeler\dava-notu.md"
mahrem mask "C:\Belgeler\dava-notu.md"
mahrem mask "C:\Belgeler\dilekce.docx"
mahrem mask "C:\Belgeler\karar.pdf"

CLI rumuzlu dosya olusturur fakat surec kapaninca geri acma tablosu silinir. Geri acilabilir akista Claude/Codex icindeki uzun omurlu MCP sunucusunu kullanin.

MCP araclari:

  • scan_file: Ham degerleri gostermeden tur/adet raporu verir.

  • mask_file: Rumuzlu dosya ve modele uygun maskeli metin olusturur.

  • restore_file: Sonucu yerelde geri acar, acik metni modele dondurmez.

  • forget_session: Bellek ici eslesme tablosunu siler.

Özel terimler için örnek kural dosyasını kopyalayın. Allowlist eşleşmeleri tam değer üzerinden yapılır. Gerçek isimleri bu yerel JSON'a yazın; sohbete yapıştırmayın.

Gelistirme

uv sync
uv run python -m unittest discover -s tests -v
node --test tests/extension.test.mjs
uv run mahrem scan "$PWD\examples\ornek-belge.txt"

Guvenlik ve sinirlar

Tehdit modeli ve bilinen sınırlar SECURITY.md dosyasındadır. Kısa hali:

  • Ag istegi ve telemetri yok.

  • Eslesmeler diske yazilmaz; MCP sureci kapaninca geri acma olanagi da kapanir.

  • Geri acilmis dosya ajan tarafindan yeniden okunmamalidir.

  • Otomatik tespit uzman kontrolunun yerine gecmez.

İsteğe bağlı gerçek Chromium testi: Playwright kurulu ortamda node tests/browser-extension.cjs. Bu test yerel eklentiyi açar; kontrol onayı, kopyalama, geri açılan dosya, oturum temizleme ve ağ isteği olmamasını sınar.

Güncelleme

Önce açık işlerinizi geri açıp tamamlayın. Yeni ZIP'i indirin; masaüstü kurulumunda Kur.bat dosyasını yeniden çalıştırıp istemciyi yeniden başlatın. Tarayıcıda yeni extension klasörünü yükleyin. Yeniden başlatma veya sekmeyi yenileme, bellekteki geri açma tablosunu siler.

Yol haritasi

  1. Yerel Turkce NER ile kisi/kurum/adres tespiti

  2. Sifreli ve kullanici parolali oturum kasasi

  3. DOCX/PDF icin bicimi koruyan donusum

  4. Tarayıcı mağazası dağıtımı ve ayrı değerlendirmeyle site entegrasyonları

  5. Farkli diller icin tespit paketleri

Katki kurallari icin CONTRIBUTING.md dosyasina bakin.

Available Tools

4 tools
forget_sessionB

Erase one in-memory placeholder mapping after the restored file is created.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden. It usefully reveals that the erased mapping is 'in-memory' (i.e., ephemeral/non-persistent) and that the operation is destructive in nature. It does not state failure behavior (e.g., whether forgetting a non-existent session errors), idempotency, or permissions, so disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler, and the operation is front-loaded. It is terse rather than bloated, though the terseness trades away clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with an output schema (so return values need no explanation), the description covers the effect and timing but omits what session_id means and any error/idempotency behavior. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter, session_id, and schema description coverage is 0%, so the description is the only place semantics could be added — yet it never mentions session_id or what value identifies the mapping. The word 'one' faintly implies a selector, but format and expectations are left entirely to the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb (Erase) and a resource (one in-memory placeholder mapping), and ties it to the restore flow. However, the key noun 'placeholder mapping' is never explained, and it is not clear how it relates to the tool's own name 'session' or to sibling tools like restore_file, so the purpose is only partially pinned down.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'after the restored file is created' gives an implicit timing condition for when to call it, which is useful sequencing guidance. But no alternatives or when-not-to-use conditions are stated, and sibling tools are never named, so the agent must infer the workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mask_fileC

Mask local TXT/MD/DOCX/PDF to plain text. Detection is partial: human review is required.

ParametersJSON Schema
NameRequiredDescriptionDefault
overwriteNo
rules_pathNo
output_pathNo
source_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the important caveat that detection is partial and human review is required — genuinely useful behavioral context. But it says nothing about permissions, overwrite behavior, side effects on the source file, or whether the operation is reversible. Only the quality caveat is covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler, caution frontloaded after the core purpose. Efficiently sized for the information it conveys.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But for a 4-parameter mutation tool with no annotations and 0% schema coverage, the description leaves the parameters and most behavioral characteristics undocumented. The partial-detection caveat is useful but insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% — none of the four parameters have descriptions. The tool's description mentions no parameters at all: not source_path, output_path, rules_path, or overwrite. With low coverage the description must compensate, and it does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (mask) and target resources (local TXT/MD/DOCX/PDF → plain text). However, 'mask' is ambiguous — it could mean redact sensitive data, obfuscate, or format-convert — and the description never clarifies which. It also doesn't distinguish from sibling scan_file or restore_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no alternatives named. The hint about partial detection is a caution, not usage guidance. An agent cannot tell from the description whether this belongs before or after scan_file, or how it relates to restore_file.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_fileC

Restore placeholders into a local output file; never returns the restored clear text.

ParametersJSON Schema
NameRequiredDescriptionDefault
overwriteNo
session_idYes
masked_pathYes
output_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose one important trait: it never returns the restored clear text, a security-relevant guarantee. However, it says nothing about permission requirements, whether overwrite is destructive to an existing file, or what happens if the session is invalid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler, and the core action is front-loaded ahead of the negative constraint. Nothing is wasted, though the sentence is arguably too terse given the surrounding gaps.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema lowers the burden on return-value explanation, but the tool still leaves essential gaps: a prerequisite session_id is undocumented, overwrite semantics on an existing file are unstated, and three of four parameters have zero coverage. For a mutating file tool with no annotations, this is insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so all four parameters (masked_path, session_id, overwrite, output_path) are undocumented in both schema and description. The phrase 'local output file' loosely hints at output_path and the non-return of clear text implies the result is file-bound, but masked_path, session_id, and overwrite remain entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (restore) and resource (placeholders into a local output file), which clearly separates it from mask_file and scan_file in the sibling set. It stops short of explicitly naming how it relates to those siblings, but the action itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no mention of prerequisites (e.g., that a prior mask_file run and matching session_id are required). Usage is only inferable from the sibling set, which is weak routing information for a four-tool family.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_fileB

Scan local TXT, Markdown, DOCX or text PDF without returning raw values. Report extraction warnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
rules_pathNo
source_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; it does disclose one genuinely useful trait — that raw values are not returned — plus that extraction warnings are reported. However it is silent on whether source files are modified, what permissions are needed, whether any session state is created (relevant given forget_session exists), and how the warnings surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the operation and the formats, with the no-raw-values constraint placed prominently. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, and the description covers formats and warning reporting. But with zero annotation coverage and zero parameter documentation, the definition leaves the role of rules_path and this tool's place in the scan/mask/restore pipeline unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters: source_path and rules_path. The description clarifies source_path only indirectly via the format list, and never mentions rules_path at all, leaving the purpose and expected content of the rules file entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scan) plus the resource and its accepted formats (local TXT, Markdown, DOCX, text PDF), and adds a distinguishing constraint (no raw values returned). It is clear what the tool does, though it never names or contrasts with the sibling tools mask_file/restore_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose scan_file over mask_file or restore_file, nor any stated prerequisite or workflow position (e.g. scan before masking). Usage can only be inferred from the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.2.0
    • First observedforget_session
    • First observedmask_file
    • First observedrestore_file
    • First observedscan_file

TDQS

B3.3/5.0

Scored across 4 tools

Disambiguation5/5

Each tool maps to a distinct stage in the masking workflow: scan for warnings, mask to produce placeholders, restore to reinsert values, and forget to clean up session state. Although scan and mask both operate on local files, their outputs and purposes are clearly differentiated.

Naming Consistency5/5

All four names follow a consistent snake_case verb_noun pattern (scan_file, mask_file, restore_file, forget_session), making the set predictable and easy to navigate.

Tool Count5/5

Four tools is well-scoped for a focused privacy-preserving file workflow; each tool earns its place without redundancy.

Completeness4/5

The set covers the core lifecycle of scan, mask, restore, and cleanup, but lacks any tool to list or inspect active sessions/placeholder mappings, which could be a minor gap for multi-file use.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Provides local anonymization of Czech legal documents by replacing sensitive entities with pseudonyms to ensure privacy during LLM interactions. It allows users to safely process documents like contracts and judgments by keeping original data offline and facilitating local deanonymization.
    5
    5
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Sanitizes text and files by removing PII, secrets, and custom patterns locally before sending to LLMs, with optional reverse-scrubbing.
    3
    105 npm
    2
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to interact with local documents (PDF, Markdown, TXT) through tools for discovery, reading, extraction, summarization, comparison, keyword extraction, search, and analysis, ensuring privacy and offline capability.
    -
  • A
    license
    A
    quality
    B
    maintenance
    Anonymizes documents via the Stript desktop app locally, ensuring personal data never enters the conversation. Provides tools to anonymize files and clipboard content, fetch results, and restore original values to local files.
    5
    181 npm
    MIT