I'm the person businesses bring in when something's broken and needs a specialist, not a generalist — reviewing, correcting, and refining AI systems at the level most consultants never touch.
I started as a scientist. Lab work at Rutgers, validated protocols, nutritional and mineral analysis — the kind of work where precision isn't optional. That precision is what got me pulled into AI training early: I wasn't just writing prompts, I was quickly promoted to reviewing and correcting the work of other researchers and engineers, because I could see what was wrong before anyone else could.
That pattern — being the one who catches what's broken — is the same reason I went to Columbia Engineering to learn machine learning, neural networks, and transformers from the inside instead of using them as a black box. It's also why small and mid-sized businesses come to me: not because they need another AI vendor, but because they've usually already hired the wrong one, and need someone who can actually diagnose what's broken and fix it.
Same instinct, applied at every scale.
NDA-safe summaries of 26 frontier AI training and evaluation projects. Hover to pause, scroll to explore.
Reviewed paired videos of animated characters — faces covered, no sound — comparing body-movement quality to train model understanding of non-verbal, gesture-based signal independent of speech or expression.
Assessed model outputs for visual grounding and spatial reasoning accuracy using rubric-based evaluation, improving model performance on spatial tasks.
Reviewed and labeled first-person video footage captured via wearable smart-glasses hardware, producing high-quality datasets for computer vision and multimodal model training.
Designed adversarial prompts to probe model safety boundaries and refusal logic, documenting vulnerabilities to strengthen guardrail robustness.
Reviewed model chain-of-thought outputs for logical coherence and reasoning accuracy, improving performance on complex multi-step reasoning tasks.
Built a custom rubric to rate and rank competing image-based model responses, enhancing LLM capability in image analysis and interpretation.
Evaluated model responses to math and physics reasoning tasks for accuracy and logic, improving logical consistency in STEM domains.
Developed rubric criteria and gold-standard example responses to calibrate evaluator judgment across a distributed team, producing standardized, reliable training datasets.
Designed complex image-based prompts to surface model failure modes, then refined responses to fix identified issues and strengthen robustness.
Classified images and prompts by criteria, evaluated responses for failures, and refined for clarity and alignment across multimodal outputs.
Composed image-based prompts, rated model responses, and refined them through reinforcement-learning-style feedback to enhance visual analysis capability.
Evaluated paired model responses, selected the stronger output, and iteratively hardened prompts to surface and correct difficult failure cases.
Conducted stress-test-level fact-checking on technical, hard-to-verify queries, producing high-fidelity training data for model factuality.
Reviewed model outputs to identify false or unnecessary refusals and assessed logical consistency, lowering the rate of over-refusal on legitimate requests.
Wrote multi-turn, entirely human-authored conversations from source material, maintaining coherence across long-form, multi-document training seeds.
Ranked paired model outputs based on quality and human-preference alignment, improving human-satisfaction scores for model outputs.
Rated model responses, wrote rubrics defining what a good response looks like, then rewrote outputs to improve one of the world's most widely used chatbots.
Replaced identifying information with fictional data, tagged identifier types, and maintained conversational flow and consistency across platforms for privacy compliance.
Evaluated and refined visual reasoning outputs for a state-of-the-art vision-language model, contributing to alignment across image-and-text tasks.
Rated competing model responses across accuracy, relevance, completeness, grammar, verbosity, and depth, providing justified preference rankings.
Evaluated chatbot answers against a rubric covering bias, harm, accuracy, detail, grammar, tone, and format, contributing to safer, higher-quality outputs.
Evaluated prompts for sensitive content, crafted well-structured responses, and cited external sources to meet high readability and accuracy standards.
Compared model-generated claims to source text, determined support level, and documented specific attributions to ensure trustworthy training data.
Wrote complex chemistry-domain prompts using trusted reference sources, then rated and ranked model responses to train safe, accurate technical chatbots.
Wrote prompts with 7+ constraints to stress-test chatbot completeness and accuracy; top performers were promoted to multi-turn evaluation.
Wrote prompts across creative and analytical domains, selected the stronger of two AI models' responses, and edited to improve quality.
Reviewed AI-assisted rewrites via a quality rubric, made manual corrections, and preserved 100% human-authenticity under a strict no-AI-tool constraint.
Reviewed and rated rewrites via a dimensionalized quality rubric, providing granular editorial feedback to ensure high-quality, accurate technical text.