Technical Case Study8 min read

Building VigilDesk: Practical RAG Copilot, Injection Defense & Human-in-the-Loop Triage

How I built a multi-tenant support triage agent and embedded copilot (Pip) using Flask and React to eliminate enterprise documentation bottlenecks with grounded safety controls.

Daniel Giovinazzo
Daniel Giovinazzo
Full-Stack Software Engineer

The Operational Bottleneck: Why Support Teams Stall#

Before I ever wrote enterprise code, I spent nearly a decade working in my family’s high-volume Italian restaurant in New Jersey, followed by years in direct field sales. If there is one lesson that operations beat into me, it is this: when people are under intense pressure, bottlenecks are rarely caused by a lack of effort—they happen because the information people need to do their jobs is fragmented or out of reach.

In enterprise IT and HR support, tier-1 specialists spend up to 40% of their day digging through disconnected Confluence spaces, outdated wikis, and dense internal PDF policy manuals. While SLA clocks tick down, agents are stuck cross-referencing conflicting documents to answer basic questions about parental leave, VPN credentials, or travel expense reimbursements.

I built VigilDesk to solve that exact friction. It is a full-stack, multi-tenant autonomous support triage agent and RAG knowledge service built with Python/Flask on the backend and React 18 (TypeScript) on the frontend. It pairs support teams with an embedded AI copilot named Pip to handle policy retrieval and draft responses, while keeping human specialists firmly in control of every customer-facing output.

The VigilDesk Unified Triage Cockpit
The VigilDesk Unified Triage Cockpit: Live SLA Ticket Queue, Interactive Workbench, and Context-Aware Pip AI Copilot.

System Design: Grounded Knowledge Retrieval without Hallucinations#

The biggest risk with deploying LLMs into customer support workflows is ungrounded hallucinations. An agent cannot guess an HR policy or invent an IT troubleshooting step.

To keep answers strictly grounded, I implemented a hybrid vector retrieval pipeline. In enterprise environments, dense vector embeddings alone often fall short when users search for exact part numbers, software version codes, or specific policy acronyms. By pairing dense pgvector embeddings with sparse lexical indexing (BM25), the retrieval engine balances semantic conceptual matching with exact keyword precision.

Every tenant document is partitioned with structural metadata preservation—retaining document titles, section headers, versions, and tenant workspace IDs to ensure strict database-level isolation between organizations.

Draft with Pip Live Demo with RAG policy retrieval
Triggering "Draft with Pip" on an active ticket: intermediate RAG retrieval traces and grounded response generation.
pgvector_hybrid_query.sql
sql
-- Multi-tenant vector retrieval with strict workspace boundary isolation
SELECT 
    d.id,
    d.content,
    d.metadata,
    1 - (d.embedding <=> :query_vector) AS cosine_similarity,
    ts_rank_cd(d.search_vector, plainto_tsquery('english', :text_query)) AS text_rank
FROM tenant_knowledge_chunks d
WHERE d.workspace_id = :workspace_id
  AND d.is_active = true
ORDER BY (
    0.7 * (1 - (d.embedding <=> :query_vector)) + 
    0.3 * ts_rank_cd(d.search_vector, plainto_tsquery('english', :text_query))
) DESC
LIMIT 5;

Defensive Thinking: Prompt Injection & Context Encapsulation#

I have played competitive chess since I was eight years old. Chess teaches you to evaluate board vulnerabilities before you make an offensive move. In software, when you build an agent that reads user-submitted support tickets, you are exposing your system to completely untrusted text.

A malicious user or a formatted payload inside a ticket can attempt a prompt injection attack—trying to override system prompts, extract API keys, or trigger unauthorized database updates. To protect against this, I implemented a dual-stage defensive boundary:

First, inbound ticket payloads run through heuristic pattern scanners to detect adversarial tokens before the model sees them. Second, untrusted text is encapsulated inside cryptographically delimited, immutable XML blocks, explicitly instructing the model that tokens within those tags cannot alter execution logic.

injection_guard.py
python
def sanitize_ticket_payload(user_input: str, tenant_context: dict) -> SanitizedPrompt:
    # 1. Check for token boundary bypass attempts
    if contains_adversarial_patterns(user_input):
        log_security_telemetry(tenant_context["workspace_id"], "injection_attempt_flagged")
        raise SecurityPolicyViolation("Input matches known prompt injection signatures.")

    # 2. Strict XML context boundary encapsulation
    safe_wrapped = (
        "<UNTRUSTED_CUSTOMER_INPUT>
"
        f"{html.escape(user_input)}
"
        "</UNTRUSTED_CUSTOMER_INPUT>"
    )
    return SanitizedPrompt(safe_wrapped, verified=True)

Human-in-the-Loop: Keeping Specialists in Control#

Autonomous tools should empower human operators, not replace their judgment. If an AI agent unilaterally issues refunds, modifies account credentials, or closes escalated tickets, it creates catastrophic liability.

In VigilDesk, I implemented a bounded reasoning loop with stateful Human-in-the-Loop (HITL) safety. When a specialist opens a ticket, the Pip copilot reviews the context, searches the policy base, and drafts a structured response with direct source citations.

The draft is loaded into an interactive workbench editor. The specialist can review the cited policy snippet, tweak the wording, and approve the send with a single keystroke. No customer communication or mutating tool execution ever occurs in the dark.

VigilDesk Observability & Audit Dashboard
Real-time observability dashboard tracking agent execution latencies, token consumption, and step-by-step traces.

Practical Takeaways: Engineering for Real Users#

Support specialists do not need flashy conversational gimmicks; they need reliable tools that save them time and prevent mistakes. That is why I also invested heavily in mobile-first ergonomics, designing a responsive master-detail layout so on-call engineers can review escalated tickets and approve drafts directly from their phones.

Building VigilDesk reinforced my core engineering philosophy: complex operational problems are solved by practical, maintainable software with clean boundaries, solid test coverage, and deep respect for the people who use it every day.

VigilDesk Mobile Triage Interface
Dedicated master-detail mobile layout designed for on-call support engineers triaging tickets with SLA timers and priority tags.

Collaboration & Engineering Credits#

Enterprise software is rarely built in isolation. VigilDesk was developed as a collaborative engineering effort alongside Jameson Wang and Wendy Gong as part of Flatiron School's AI & Data Science program.

Delivering a dependable AI support platform required continuous cross-functional teamwork—from architecting multi-tenant vector retrieval and bounded reasoning loops to establishing prompt injection defenses and fine-tuning specialist triage ergonomics.

Special thanks to Jameson and Wendy for their dedication and engineering collaboration throughout the architecture and validation of VigilDesk.

Project Contributors & Engineering Credits

Jameson Wang
Jameson Wang
Full-Stack Software Engineer
Wendy Gong
Wendy Gong
AI Engineer / Product
Topics:Autonomous AI AgentRAG / Vector EnginePrompt Injection DefenseTypeScriptPython / FlaskFull-Stack Development
Let's Build Together

Have an engineering challenge or role in mind?

I'm open to full-stack software engineering and AI engineering opportunities. Reach out to discuss technical architecture, system design, or team collaboration.