Projects

Gates, harnesses, and delivered systems.

Two bodies of work: open-source evaluation and governance harnesses published on GitHub — runnable, tested, and CI-checked on synthetic data — and professional engagements delivered inside employer and client environments. The two are kept separate on purpose.

Open-source evaluation & governance harnesses

Nothing ships until it passes a gate — and the gate itself must be earned.

Seven runnable, tested, CI-checked projects published on GitHub. Every dataset is synthetic, every repository documents its own limits, and none of these are employer or client production deployments.

[00]

jbisaccia-9/rag-gate

Retrieval quality gate for RAG systems — the index only serves once retrieval is good enough.

Gate
Serves only at recall@3 ≥ 0.90.
Evidence
Baseline caught at 0.83, then fixed to 1.00 on the current small synthetic set.
Limits
Synthetic corpus; README notes benchmark saturation and the limits of a small labeled set.
RAGRetrieval evalCI gate

[01]

jbisaccia-9/kappa-gate

Calibration harness for LLM-as-judge evaluation using Cohen's kappa against human labels.

Gate
Requires kappa ≥ 0.70 and agreement ≥ 0.85 before a judge is trusted.
Evidence
The mock judge is refused; a recorded claude-opus-5 run passes.
Limits
Dimension-level limitations documented in the README; small labeled sample.
LLM-as-judgeCohen's kappaEvaluation

[02]

jbisaccia-9/perm-gate

Prompt-layer guardrails versus permission-layer enforcement, measured side by side.

Gate
Access decisions must be enforced below the prompt, not inside it.
Evidence
On the synthetic test set, prompt mode leaked 4/5; permission mode leaked 0/5.
Limits
The assistant under test is deliberately not an LLM — the README explains why that isolates the variable.
AuthorizationGuardrailsSecurity

[03]

jbisaccia-9/phi-gate

Regex-tier redaction gate for PHI-shaped identifiers in healthcare-adjacent text.

Gate
Redaction must clear the recall and precision floor before text moves downstream.
Evidence
Recall 1.00 and precision 0.95 on the current synthetic corpus.
Limits
Free-text names and addresses are explicitly out of scope for the regex tier.
PHIRedactionHealthcare

[04]

jbisaccia-9/roi-gate

A deliberately conservative model for AI adoption ROI, built to resist vendor-deck math.

Gate
CI refuses aggressive assumptions before a number can be published.
Evidence
On the synthetic example, the conservative model reports $24,048/yr against $517,704/yr under vendor-deck assumptions.
Limits
Illustrative synthetic inputs — a modeling tool, not a forecast of any organization's results.
ROI modelingAssumption gatingCI

[05]

jbisaccia-9/target-gate

Twice-monthly provider-targeting pipeline where list delivery is blocked until quality gates pass.

Gate
Identifier, freshness, dedupe, coverage, and brief-grounding gates must all pass before delivery.
Evidence
Synthetic fixtures committed to the repo; production-shaped adapters for Azure Functions, Foundry, and Graph.
Limits
Fixtures are synthetic; adapters are production-shaped rather than a deployed production system.
PipelinesData qualityAzure

[06]

jbisaccia-9/trade-gate

Order-validation guardrail design study — how a validation layer refuses malformed intent.

Gate
Orders must clear validation rules before they are ever considered valid.
Evidence
Synthetic fixtures with tested, CI-checked validation rules.
Limits
Educational design study. Not a trading system and not financial advice.
ValidationGuardrailsEducational

Professional engagements

Delivered inside employer and client environments.

~/projects/hipaa-clinical-ai-function▍

01 · Behavior Frontiers · Jul 2026 – Present

Enterprise AI function for a national behavioral health network

Outcome

Established the organization's AI function and the training and change-management program that lets distributed teams adopt AI tools responsibly. Work is in progress; outcome metrics are not yet published.

What it enabled

Serving as the founding AI engineering resource: building HIPAA-compliant AI infrastructure, RAG pipelines, agentic workflows, and LLM-powered automation for clinical and operational teams.

Context

A national network of autism and behavioral health centers had no in-house AI engineering capability and no compliant path to apply LLMs to clinical and operational work.

Architecture

Retrieval pipelines over internal documentation, agentic workflows for repeatable operational tasks, and LLM automation integrated with existing enterprise systems.

Security

HIPAA-compliant infrastructure design, PHI-aware data handling, and secure enterprise AI architecture patterns for a regulated clinical environment.

Governance

Governance-oriented implementation priorities set with clinical, operations, and department stakeholders; LLM evaluation and hallucination-mitigation practices applied to deployed workflows.

Technologies

PythonLangChainRAGAgentic workflowsAPI integration
~/projects/solar-rag-chatbot▍

02 · Capital Energy · 2024 – Jul 2026

Customer-facing RAG chatbot with agentic logic

Outcome

Reduced response times by 40% and increased self-service adoption for inbound queries.

What it enabled

Designed and deployed a customer-facing RAG chatbot with agentic logic that answers inbound solar queries directly and hands off when human help is needed.

Context

Inbound solar inquiries arrived faster than the team could answer them, slowing response times and pushing routine questions onto sales staff.

Architecture

Retrieval over the company's product and process documentation, agentic routing for multi-step questions, and integration with existing customer channels.

Security

Scoped credentials for integrated systems and constrained retrieval sources to approved company content.

Governance

Prompt iteration guided by output review and response-quality evaluation before broader rollout.

Technologies

RAGPrompt engineeringAgentic logicAPI integration
~/projects/lead-reactivation-agent▍

03 · Capital Energy · 2024 – Jul 2026

Outbound lead reactivation agent

Outcome

Automated prospect engagement at scale and accelerated sales pipeline growth.

What it enabled

Built an outbound reactivation agent using Voiceflow, Twilio, and ElevenLabs to automate prospect engagement and route interested leads back to sales.

Context

A large backlog of dormant prospects sat untouched because manual outreach did not scale with the sales team's capacity.

Architecture

Conversation flows in Voiceflow, telephony and messaging via Twilio, synthesized voice via ElevenLabs, with outcomes written back to the CRM.

Security

Scoped API credentials per integrated service and controlled contact lists for outreach.

Governance

Human handoff for qualified conversations and review of agent transcripts to tune behavior.

Technologies

VoiceflowTwilioElevenLabsCRM administration
~/projects/crm-migration-automation▍

04 · Capital Energy · Technical project management

CRM migration and operational automation program

Outcome

Reduced manual operational work by 40% and setup time by 30%.

What it enabled

Led the CRM migration to Core 365 as technical project manager — data migration, workflow redesign, and system integration — and designed Make and Zapier automation pipelines across operations.

Context

Fragmented systems and manual handoffs made operational work slow to run and slow to set up for new campaigns and teams.

Architecture

Migrated CRM data model and redesigned workflows, with Make and Zapier pipelines connecting CRM, communications, and internal tooling.

Security

Controlled data migration with scoped access during cutover and per-connection credential management.

Governance

Staged migration plan with stakeholder sign-off, workflow documentation, and post-cutover support.

Technologies

Core 365MakeZapierCRM administrationAPI integration
~/projects/frontier-model-training▍

05 · Handshake AI · Outlier AI · Mercor · 2024 – Present

Model training, evaluation, and RLHF contract work

Outcome

Directly informs how I design evaluation and hallucination-mitigation practices for enterprise deployments.

What it enabled

Ongoing contract work performing expert data annotation and dataset curation for LLM training pipelines, evaluating outputs against reward metrics, and contributing to RLHF and preference-data workflows.

Context

Frontier AI platforms need expert human judgment to curate training data and evaluate model behavior on technical and conversational tasks.

Architecture

Platform-provided annotation and evaluation environments with rubric-based scoring and preference comparison tasks.

Security

Work performed under each platform's confidentiality and data-handling requirements; project specifics are not disclosed.

Governance

Rubric-driven scoring focused on response quality, alignment, consistency, and hallucination mitigation.

Technologies

LLM evaluationRLHFPreference dataPrompt engineering

Representative capabilities

What I can build for an engagement.

Capability areas rather than delivered case studies — scoped to a client’s environment during discovery.

[00]

RAG and retrieval infrastructure

Ingestion, chunking, and retrieval over private corpora with citation-first synthesis and access-aware sourcing.

[01]

Agentic workflow development

Tool-using agents scoped to defined business processes, with human handoff at the points that need judgment.

[02]

LLM evaluation & hallucination mitigation

Rubric and reward-metric evaluation of model outputs, applied before rollout and revisited as prompts and models change.

[03]

Governed, compliant AI architecture

Secure enterprise AI patterns for healthcare and other regulated industries, including HIPAA-aware infrastructure design.

[04]

Automation & systems integration

Make, Zapier, and API-level integration across CRM, communications, and internal tooling.

[05]

Adoption, training & change management

Stakeholder engagement and training enablement so distributed teams actually use what gets built.

All seven gate projects live on GitHub

Clone, run the tests, and read the limits each repository documents. Client and employer work stays private.

github.com/jbisaccia-9