Skip to content

Projects

Projects

Robotics & Embodied AI

Natural-Language Task Planner for ABB GoFa Robot

Azure OpenAI
MCP
RAPID
RWS
Task Planning
ABB GoFa

Built an LLM planner that turns natural-language instructions into consecutive and executable operations for an ABB GoFa robot.

  • Decomposes instructions into multi-step task lists and coordinates parallel robot operations with execution-status tracking.
  • Built an MCP server over Robot Web Services and RAPID for equipment queries, motion, picking, placing, and tool-call feedback.
  • Added execution-failure handling and short-term context so the planner can recover and continue the workflow.
Placeholder for the ABB GoFa natural-language task planner

Voice-Guided Embodied AI Robot for Universal Grasping

LangGraph
VLA
FastSAM
GraspNet
MCP
UR Robot

Created a vision-language-action agent for voice-guided UR robot grasping and tool-delivery tasks for the 12th NVIDIA Sky Hackathon.

  • Converted spoken requests into real-time task plans, tool selections, and robot actions through a LangGraph agent.
  • Combined FastSAM and vision-language models for scene understanding with Strategy of Mind prompting and GraspNet for grasp planning.
  • Exposed UR robot actions through MCP tools for collision-free pickup and delivery.
  • Our project received the Excellence Award (National Second Place).
Placeholder for the voice-guided embodied AI grasping robot

Autonomous Mobile Crane Robot for Object Transportation

STM32
C
CAN Bus
Path Planning
Mechanical

Built a competition robot that autonomously transported and placed objects along designated routes.

  • Designed the mechanical structure and integrated sensing, actuation, and control.
  • Coordinated DC motors, encoder motors, servos, and robotic arms through an STM32 controller and CAN bus.
  • Added a dynamic shortest-path algorithm and moved from one to three motors, cutting operation time from 4 min 32 s to 3 min 15 s.
  • Our project received the National First Place Award.
Placeholder for the STM32 autonomous mobile crane robot

LLM-Based AI Agents

Deep Agents: Harness-Engineered Web Automation

Deep Agents
LangGraph
React
Vite
Skills
MCP

Built a controlled agent harness that turns visual references, wireframes, and content into responsive web pages with isolated browser previews.

  • Orchestrates visual analysis, design, HTML implementation, validation, review, and preview publication as one observable workflow.
  • Delegates bounded work to design, implementation, and quality-review subagents, streaming plans, reasoning, tool calls, and approvals.
  • Relies on thread-scoped state, Skills, filesystem isolation, retries, call budgets, and deterministic HTML quality gates for verifiable artifacts.
Placeholder for the Deep Agents web automation interface

LLM Workstation for Flexible Manufacturing

LLM Agents
RAG
Workflows
WebSocket
Digital Assets

Developed a configurable workstation that binds language models, manufacturing knowledge, digital assets, and production-line operations into reusable agent workflows.

  • Configures agents with roles, goals, models, knowledge bases, digital assets, and assigned tasks.
  • Composes agents and Crews in a visual workflow with retrieval, model-acquisition, parsing, and equipment-control tools.
  • Connects to production-line devices over WebSocket for control, database queries, data analysis, and monitoring.
Placeholder for the flexible-manufacturing LLM workstation

AI Agents Town: Interactive Multi-Agent Virtual World

LangChain
Multi-Agent
RAG
Phaser
React

Created a browser-based town where configurable AI characters keep context and interact with users and one another.

  • Lets users define NPC names, demographics, and personalities, and manage relationships between characters.
  • Integrates LangChain agents, contextual memory, and retrieval-augmented reasoning into a Phaser environment.
  • Supports multiple maps with import/export of reusable world and character configurations.
Placeholder for the interactive AI Agents Town

Full-Stack Web Platforms

Bilibili Audience Feedback Intelligence Platform

LLM Agents
Streaming
Data Viz
Bilibili
Evidence

Built an AI review workspace that turns real Bilibili danmaku and comments into evidence-backed insights for creators.

  • Accepts a video URL, BV ID, or b23.tv short link, retrieves the real feedback, and shows the data distribution before analysis.
  • Streams the analysis to extract high-frequency topics, sentiment, valuable feedback, and actionable content suggestions.
  • Links every insight back to the original comments or danmaku, with views for user levels, danmaku density, sentiment timelines, and priority feedback.
Placeholder for the Bilibili audience feedback analysis workspace

The Three Stooges: Research Collaboration Platform

LLM Agents
Research
Literature
Community
Full-Stack

Built a research website for organizing literature, exchanging technical ideas, and discussing questions with specialized LLM agents.

  • Lets researchers curate their top 100 papers and organize discussions by technical board, with viewpoints, comments, and links back to source material.
  • Includes configurable research agents and a fine-tuned digital-twin assistant for deeper academic question answering.
Placeholder for The Three Stooges research collaboration platform

WhatsApp AI Property-Service Agent

WhatsApp
Twilio
Webhooks
WebSocket
AI Agent
Work Orders

Built an AI-assisted messaging and work-order system that answers rental customers and helps staff manage service requests.

  • Receives WhatsApp messages through a Twilio webhook and synchronizes them with the backend in real time.
  • Gives staff a WebSocket workspace to review the conversation and agent reply, and turns each request into a traceable work order with status, priority, property, and approval details.
Placeholder for the WhatsApp property-service agent workspace

Multimodal LLM Chrome Extension

Chrome
LLM
VLM
Web Search
Knowledge
RAG

Developed an in-browser AI assistant that combines webpage context, visual understanding, search, and personal knowledge in one conversation panel.

  • Answers questions against the active webpage without moving content into a separate chat application.
  • Uses a vision-language model for image recognition and visual question answering, plus web search and knowledge-base access.
Placeholder for the multimodal LLM Chrome extension

©Ding Gao 2026. All rights reserved.

Site Views: Loading....( Statistics powered by不蒜子统计)