| title | TranscriptAI |
|---|---|
| emoji | 🎙️ |
| colorFrom | pink |
| colorTo | purple |
| sdk | docker |
| app_port | 7860 |
| pinned | false |
Multilingual Meeting Intelligence · Japanese · Hindi · English · Mixed
Turns any meeting transcript or audio file into structured business intelligence in ~3 seconds. The only meeting AI that understands what Japanese and Indian business partners actually mean.
Generic meeting summarisers extract what was said. They miss what was meant.
| What was said | Generic AI | TranscriptAI |
|---|---|---|
| 検討いたします | "Action: We will consider it" | ⚠ Soft rejection — 72% confidence. Follow up explicitly. |
| 難しいかもしれません | "It may be difficult" — neutral | 🚨 HIGH rejection signal — 90% confidence. Deal at risk. |
| 前向きに検討 | "Positive consideration" | ⚠ Uncertain — outcome not guaranteed (55%) |
| 承知いたしました | "Acknowledged" | 🏯 はい trap — understanding, NOT approval |
| 善処します | "Action: We will handle it" | 🚨 Classic nemawashi dodge — no real commitment |
| パートナーシップは継続しないことを決定しました | "Decision made" | ⛔ CRITICAL — Explicit contract termination. Irrevocable. |
| dekhte hain | "We'll see" | ⚠ Hindi deferral — classic avoidance signal |
| kal pakka | "Definitely tomorrow" | ⚠ Fake urgency — indefinite future in disguise |
Japanese enterprise also mandates APPI compliance — raw meeting data cannot be sent to foreign cloud LLMs. Most tools fail this requirement by design. TranscriptAI masks all PII locally before any LLM call.
Input: Transcript (JP · HI · EN · Mixed) or Audio (MP3/MP4/WAV/M4A)
Output: Structured business intelligence in ~3 seconds
- 8-state meeting outcome verdict — 🟢 Approved · 🔴 Rejected · 🔵 Conditional · 🟣 Deferred · 🟡 Pending · ⚪ Informational · 🟠 At Risk · ⚫ Unclear
- 60+ rejection patterns across 3 tiers — CRITICAL (explicit termination) → HIGH (performance-failure framing) → MEDIUM/LOW (soft hedging)
- 20 JP soft rejection patterns — nemawashi, 難しいですね, ぜひ検討, 対応しかねます and more with confidence scores
- 8 Hindi indirect communication patterns — देखते हैं, थोड़ा मुश्किल, kuch na kuch ho jayega and more
- Keigo formality detection — MeCab morphological analysis, not word-level guessing
- はい trap detection — 承知しました / 承知 flagged as understanding, not approval
- Approval gate detection — "board must approve", 稟議が必要です — deal not done yet
- Local PII masking before any data leaves the server — 500+ Japanese surnames, phones, emails
- Bidirectional PIIMask —
[NAME_1]→Tanakaafter analysis, never sent to the LLM - Fully local mode via Ollama — zero cloud exposure when required
- Hallucination guard — rule-based token overlap, LLM never validates its own output
- Register-based sentiment — scores how a speaker treats the other party, not word valence (professional apologies = neutral, not negative)
- Cross-script speaker normalization — 田中 ↔ Tanaka ↔ Director resolved to same identity
- Meeting health score — 0–100 across sentiment, action clarity, communication risk, AI confidence. Capped at 35 for HIGH risk, 22 for CRITICAL/termination
- 議事録 — Japanese formal business minutes in standard enterprise structure
- PPTX deck — 6-slide presentation with said-vs-meant, risk watch, decisions, next steps
- Cultural insights — nemawashi risk, 稟議 approval status, keigo level breakdown
- Markdown / JSON / TXT — for downstream workflows
- Evaluation page — run 3 bilingual ground-truth test cases on demand, view scores
- MLflow integration — experiment tracking at
http://127.0.0.1:5000, auto-logs per run - JSONL audit log — append-only observability with schema drift detection
| Version | Key change | Score |
|---|---|---|
| v1 | Hard exact matching, English-only | 22–30% |
| v2 | Fuzzy speaker names, TF-IDF similarity | ~45% |
| v3 | MeCab keigo override, bilingual ground truth | ~60% |
| v4 | Hallucination guard, nemawashi patterns, APPI masking | 75–85% |
| v5 (live) | 2-key rotation, vector cache, bypass_cache eval fix | 93.8% |
Every accuracy improvement was driven by evaluation metric failures traced through the pipeline — not intuition. When action F1 was 0.4 at v2, tracing revealed the LLM was extracting "Director" (role title) instead of "Tanaka" (first name) as the action item owner. One prompt rule fixed it; F1 jumped to 0.87.
| Test Case | Overall | ROUGE-1 | Action F1 | Sentiment |
|---|---|---|---|---|
| Sales call · JA/EN mixed | 94.5% | 0.694 | 1.0 | 1.0 |
| Internal meeting · Japanese | 93.8% | 0.703 | 1.0 | 1.0 |
| Client conflict · EN/JA | 93.8% | 0.703 | 1.0 | 1.0 |
The prompt pipeline was fully optimized in v3.2 to minimize Groq free-tier consumption:
| Stage | Before | After | Saved |
|---|---|---|---|
_GROUNDING_RULES |
424 tokens | 94 tokens | -330 (-87%) |
| Rules block | 291 tokens | 110 tokens | -181 (-62%) |
| Schema block | 280 tokens | 55 tokens | -225 (-80%) |
| Sandwich repeat | 75 tokens | 0 tokens | -75 (-100%) |
| Per-request total | 2,728 tokens | 1,521 tokens | -1,207 (-44%) |
30 users/day: 45,630 tokens (46% of 100K free tier limit). Supports ~65 users/day.
Additional optimizations active:
- Model routing — short EN-only transcripts use
llama-3.1-8b-instant(separate quota bucket) - Dynamic schema —
japan_insightsblock only included for JP/mixed transcripts - Transcript truncation — 1,200 word cap (first 60% + last 40%) for very long meetings
- Reduced
max_tokens— capped at 550–1,100 depending on transcript length
Input transcript / audio
│
▼
1 Vector cache check utils/vector_cache.py ChromaDB cosine similarity (≥95% → instant return)
2 MD5 exact cache utils/cache.py Hash match → return in <1ms
3 PII masking transcription/pii_masker.py APPI — masks ALL PII before LLM sees text
4 LLM analysis analysis/analyzer.py Groq 70B → 8B → Ollama → Mock
5 PII restoration transcription/pii_masker.py Restores [NAME_1] → Tanaka BEFORE normalization
6 Speaker normalization transcription/speaker_normalizer.py 田中 ↔ Tanaka ↔ Director → unified
7 MeCab keigo override analysis/japanese_tokenizer.py Morpheme-level formality, overrides LLM guess
8 Code-switch count utils/evaluator.py Rule-based Unicode range detection
9 Hallucination guard analysis/hallucination_guard.py Token overlap + semantic similarity
10 Rejection + outcome analysis/soft_rejection_detector.py + deal_outcome_detector.py
11 Cache + log utils/vector_cache.py + utils/logger.py
Critical ordering: PII masked before step 4 (LLM). PII restored before step 6 (normalization). Reversing either order breaks the pipeline.
| Layer | Choice | Why Not the Alternative |
|---|---|---|
| LLM inference | Groq (llama-3.3-70b) | 10–20× faster than GPU APIs. Free tier. JSON mode. |
| Local fallback | Ollama (qwen3:8b) | Zero cloud exposure for strict APPI cases |
| Vector DB | ChromaDB | Free, local, HF Spaces compatible, APPI compliant |
| Japanese NLP | MeCab + IPADIC | Morpheme-level auxiliary verb detection — keigo is invisible to word-level tokenizers |
| Web framework | FastAPI | Native async, asyncio.to_thread(), auto Swagger |
| Frontend | Alpine.js + Jinja2 | No bundler, no build step, HF Spaces compatible |
| ML tracking | MLflow | Free, local SQLite, APPI compliant |
| PPTX | python-pptx | Full slide control |
| Audio | Groq Whisper | Free tier, fast, multilingual |
git clone https://github.com/aiKunalBisht/Transcript-ai.git
cd Transcript-ai
pip install -r requirements.txt
# Required
export GROQ_API_KEY=your_key_here
# Optional — enables 2-key round-robin (use keys from DIFFERENT accounts for separate quotas)
export GROQ_API_KEY_2=your_second_key_here
# Start
uvicorn main:app --reload --port 7860
# Open http://localhost:7860
# API docs at http://localhost:7860/docsFully local — zero cloud exposure:
ollama pull qwen3:8b
# No config needed — app auto-detects Ollama when no Groq key setmain.py FastAPI server — routes, module loading, speaker label detection
analysis/
analyzer.py LLM orchestration — provider chain, prompt, token optimization
soft_rejection_detector.py 3-tier rejection detection — CRITICAL / HIGH / MEDIUM / LOW
deal_outcome_detector.py 8-state meeting outcome verdict — Approved through Unclear
hallucination_guard.py Rule-based token overlap verification
conversation_dynamics.py Topic stalls, senior silence pivots, closing summarizer
japanese_tokenizer.py MeCab morphological keigo detection
english_analyzer.py 40+ EN hedging and commitment-strength patterns
hindi_analyzer.py 8-category Hindi/Hinglish indirect communication patterns
semantic_validator.py Sentence-transformer semantic similarity
agents/
gijiroku_formatter.py 議事録 Japanese business minutes generator
cultural_insights_formatter.py Cultural context export
slide_architect.py PPTX slide plan — deterministic + LLM-narrative split
exporters/
pptx_builder.py python-pptx builder (670 lines)
transcription/
pii_masker.py APPI-compliant PII masking — 500+ JP surnames
audio_processor.py Groq Whisper audio transcription
speaker_normalizer.py Cross-script identity resolution
rags/
meeting_store.py ChromaDB meeting store for historical retrieval
rag_retriever.py Semantic retrieval over past meetings
utils/
html_renderer.py Results HTML — health score, outcome badge, 5-tab layout
evaluator.py Ground-truth scoring — ROUGE, F1, sentiment, MLflow
vector_cache.py ChromaDB semantic cache
logger.py JSONL audit log with drift detection
templates/
base.html Layout, sidebar, nav, Alpine.js reactive state
index.html Main analysis page
export.html Export page — PPTX, 議事録, MD, JSON, TXT
evaluate.html Evaluation page — idle until run clicked
tests/
test_core.py 21 pytest tests
test_data.py 3 bilingual ground-truth test cases (TC001–TC003)
POST /analyze-text transcript: str, language: str|null, mask_pii: bool
POST /transcribe file: UploadFile (audio or text)
POST /export/pptx result: dict → PPTX binary
POST /export/gijiroku result: dict → markdown string
POST /export/cultural-insights
POST /export/markdown
POST /export/json
POST /export/txt
GET /evaluate Evaluation page (idle — run on demand)
POST /evaluate/run Runs 3 ground-truth cases, returns scored HTML
GET /health Module availability report| Metric | Score |
|---|---|
| Performance | 94 |
| Accessibility | 100 |
| CLS (Cumulative Layout Shift) | 0.000 |
| Speed Index | 1.2s |
Built by Kunal Bisht AI/ML Engineer · LLM Pipelines · RAG · Multilingual NLP · FastAPI Pithoragarh, Uttarakhand, India · Open to Remote / Relocation