Skip to content
View Geraldxm's full-sized avatar

Block or report Geraldxm

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. opd-test-time-scaling opd-test-time-scaling Public

    Code and data for Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

    Python 3

  2. math-eval math-eval Public

    Reproducible math reasoning eval with resumable generation and separate scoring. Swap parsers without re-calling models.

    Python 6

  3. math-vault math-vault Public

    Curated, traceable snapshots of public math reasoning datasets. Canonical JSONL format with provenance tracking — plug directly into math-eval.

    Python 7

  4. resistzzz/Co-rewarding resistzzz/Co-rewarding Public

    [ICLR2026] "Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models"

    Python 30 4

  5. DIYgod/RSSHub DIYgod/RSSHub Public

    🧡 Everything is RSSible

    TypeScript 46.4k 10.2k

  6. JLU-CS-Courses JLU-CS-Courses Public

    吉林大学计算机科学与技术专业,我的课程资料、复习题和笔记等。

    HTML 242 38