Skip to content
View shushuyang231's full-sized avatar
  • Shanghai Jiao Tong University
  • Shanghai ,China
  • 17:45 (UTC +08:00)

Block or report shushuyang231

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
shushuyang231/README.md

Sun Shengyao (孙圣尧)

Incoming second-year undergraduate at SJTU SPEIT | LLM & software-agent evaluation · interested in embodied intelligence & robotics


About

Undergraduate entering my second year at Shanghai Jiao Tong University, enrolled in the Paris Elite Institute of Technology (SPEIT). My prior research experience is in LLM evaluation: building evaluation benchmarks for LLM agents and studying structured-output robustness. I am now actively expanding toward embodied intelligence and robotics, while continuing to explore AI evaluation.

Areas of Interest

  • Large language models(大语言模型)
  • Embodied intelligence & robotics(具身智能与机器人学)
  • Mechanical systems(机械)

Publications

  • FrameShift-CAD: Executable Coordinate-Frame Interventions for Diagnosing Text-to-CAD Generation — Research Square preprint, 2026. DOI
  • Did It Happen? Counterfactual Evaluation of LLM Agent Recovery from Ambiguous Tool Outcomes — Research Square preprint, 2026. DOI
  • Mutation Testing of Task-Scoped State Oracles in Software-Agent Benchmarks: A Cross-Benchmark Empirical Study — Research Square preprint, 2026. DOI
  • Testing JSON Schema Instruction Artifacts: Distributional Robustness under Validation-Equivalent Serialization and JSON Mode — Research Square preprint, 2026. DOI

Skills

  • Python — basic working proficiency(具备基础实用能力)
  • AI-native workflows — comfortable driving AI-assisted research and engineering workflows(AI 原生工作流)

Projects

Research:

Other:

Education

B.Eng. — Shanghai Jiao Tong University, Paris Elite Institute of Technology (SPEIT) 2025 – Present · Shanghai, China

Connect

Pinned Loading

  1. schema-order-robustness schema-order-robustness Public

    Reproducible study of JSON Schema serialization-order robustness in black-box LLM generation

    Python

  2. ambiguous-tool-outcomes-benchmark ambiguous-tool-outcomes-benchmark Public

    Counterfactual benchmark for LLM agent recovery from ambiguous tool outcomes

    Python

  3. side-effect-calibration-study side-effect-calibration-study Public

    Key-free replication artifact for mutation testing of task-scoped state oracles in software-agent benchmarks

    Python

  4. snowbound-school-mystery snowbound-school-mystery Public

    《雪闭校园》Ren'Py 悬疑视觉小说源码

    Ren'Py