Post-training
Reinforcement learning and verifiable feedback for language, scientific, and agentic systems.
system.profile / online
I develop methods for large language models, reinforcement learning, and multimodal scientific foundation models. I am a Computer Science PhD student at UC Irvine, with an emphasis on reliable evaluation.
Curated commands only—this terminal does not execute
code. Try help.
trace://research-loopillustrative · not live telemetry
Reinforcement learning and verifiable feedback for language, scientific, and agentic systems.
Turning large general models into smaller, efficient experts without losing what matters.
Foundation models and evaluations grounded in genomics, biology, diagrams, and geometry.
Quick route
stream.01 / latest
Accepted work, research roles, and recent milestones—without the notification noise.
LLM Research Scientist Intern working on verifiable evaluation and post-training for scientific diagrams.
Differentiable combinatorial optimization for causal variant discovery in the non-coding genome.
Procedurally generated tasks for long-horizon video reasoning.
Our oral paper for the BioLaySumm 2025 shared task combines section-wise retrieval, LLM generation, and reinforcement learning for biomedical lay summaries.
Our work on L2 normalization and geodesic distance in high-dimensional single-cell visualization received the award at the 15th ACM Conference on Bioinformatics, Computational Biology, and Health Informatics (ACM BCB 2024).
workspace.02 / current
Methods and systems work, organized by the scientific question rather than the application domain.
R/01
Project lead · model distillation, evaluation, and systems
Toward sub-million-parameter expert models distilled from large genomic language models, with controlled evaluation across classification and base-resolution tasks.
R/02
Project lead · model, data, training, and evaluation
A reasoning-grounded, multimodal foundation model for cell-type-conditioned DNA generation, editing, and pair prediction. I lead the model, data, and evaluation effort.
R/03
LLM Research Scientist Intern · diagram.ai
Semantic verifiers turn structured diagram feedback into rewards for reinforcement-learning post-training and iterative repair. The public materials emphasize evaluation design and reproducible experiments; unreleased results remain high level.
R/04
Independent study · controlled empirical research
Controlled multi-domain studies across NLP and single-cell models, with multi-seed baselines, negative controls, and reproducible cluster-scale experiments.
R/05
LLM agents · knowledge distillation · on-policy training
Studying stable knowledge transfer across multi-turn agent trajectories, supported by multi-GPU training and evaluation infrastructure built with FSDP, vLLM, Ray, and SLURM.
R/06
Technical report · distributed ML systems
Hardware-aware, latency-predictable differentiable search for faster configuration and convergence of distributed machine-learning pipeline parallelism.
R/07
Public preprint · interpretable computational biology
Interpretable multi-task learning for shared epigenetic regulation across autoimmune diseases, with site–gene–pathway structure built into the analysis.
R/08
Open benchmark contributions · Kart and Minecraft proposals
I designed long-horizon video-understanding proposals for race-telemetry reconstruction in SuperTuxKart and action-ledger reconstruction in Minecraft, with seeded generators, machine-exact ground truth, deterministic graders, and anti-shortcut calibration.
Kart issue #73 ↗ Minecraft issue #74 ↗ Kart branch ↗ Minecraft branch ↗
method.lab / transfer
Two inspectable method notes: the manuscript-derived OmegaGenome loss anatomy and established on-policy curricula for multi-turn agents. Unreleased results are intentionally omitted.
archive.03 / selected
Selected peer-reviewed work across genomic discovery, video reasoning, language models, and single-cell geometry.
Efficient differentiable search for causal variants across molecular modalities and tissues.
Read paper
Section-wise retrieval, LLM generation, and reinforcement learning for accessible biomedical summaries.
Read paper
Project cells to the L2 hypersphere, compare them by angular distance, then use a spherical affinity inside SNE.
Transformer-based reinforcement learning for oracle-guided molecular de novo design.
Read paperAdapts language models for accessible biomedical communication, emphasizing readability and factual quality.
paper.controls / inspect
Move a published optimization parameter and inspect an exact reported ablation. Illustrative quantities remain separate from experimental measurements.
EARLIER / PUBLIC RECORD
Author PDFs, full author lists, and figure notes Open the academic publication section →
paper.lab / interactive reconstruction
An interactive reading of our ACM BCB 2024 paper. Follow the same feature directions from a flat simplex to a curved hypersphere, then inspect how angular distance becomes a probability neighborhood.
ACM BCB 2024 · ACM SIGBio Best Paper Award
The method projects each cell to the unit hypersphere, measures angular distance, transforms angles into spherical affinities, and normalizes those affinities before optimizing the low-dimensional embedding. The interaction below makes that causal chain inspectable.
Official PDF ↗ DOI record ↗ Method figure ↗ Embedding figure ↗ Evaluation figure ↗
The geometry control stays in view through this method trace; open Deeper math for κ, affinities, and the numbered path to the low-dimensional neighborhood.
systems.04 / field log
I move between objectives, datasets, distributed systems, product interfaces, and evaluation infrastructure.
research / scientific ML
Developed antibody-binding ΔΔG models with binding-ddg-predictor, a CarbonDesign encoder, and GearBind; created leakage-resistant complex-level SKEMPIv2 splits and balanced alanine/non-alanine sampling, improving Pearson correlation by ~10% and Spearman by ~5%.
product systems
Built a React/TypeScript rate-card workflow for VMware Cloud on AWS with Java, API Gateway, and Lambda integrations, including CSV validation and pricing-error previews; added Athena/S3/DynamoDB anomaly detection with CloudWatch alerts for usage and subscription anomalies.
open-source ML systems
Built and upstreamed a 600+ line AutoX/Python/OpenMLDB SQL pipeline that generated time-series and statistical features, then selected top features through adversarial validation, GRN, or reinforcement learning; presented the system to the open-source community.
Contribution ↗ 4Paradigm ↗ OpenMLDB ↗ Meetup talk ↗ Code Camp talk ↗multimodal learning
Developed multimodal target detection that consumes images and two audio channels to predict behavior and distance, integrating MiDaS zero-shot monocular depth estimation; also investigated multimodal neural architecture search.
Shanghai AI Laboratory ↗model efficiency / open source
Implemented Cross-Layer Equalization for data-free FP32→INT8 quantization in Intel Neural Compressor; studied NVIDIA Triton and AI Model Efficiency Tool, presented their designs to ~100 colleagues, and helped build C++ multi-framework inference tooling for CPU/GPU deployment.
Intel Neural Compressor ↗distributed medical AI
Built Horovod multi-node, multi-GPU 3D U-Net training with NVIDIA Clara, OpenMPI, and NCCL2, reaching 2.5× speedup on four GPUs across two nodes; benchmarked configurations and added Java backend support for distributed medical-imaging jobs.
Project repository ↗ Shukun Technology ↗ Horovod ↗ Technical talk ↗working stack
TEACHING
Teaching Assistant for UCI ICS 6B (Boolean Logic & Discrete Structures, Winter 2025), UCI ICS 6D (Discrete Mathematics, Spring 2025), and the University of Michigan–Shanghai Jiao Tong University Joint Institute VE370 (Computer Organization, Fall 2021).
SELECTED HONORS
SYSTEMS MODE
Python, C/C++, Java, TypeScript, SQL, shell, CUDA, PyTorch, TensorFlow, Horovod, distributed training, evaluation, and benchmark design.
playground.05 / shipped
Browser worlds, interactive mathematical essays, and tools for thinking with models.
live multiplayer / browser-native
BUILD/01
A solo-built, no-install browser RPG with six game modes, including real-time multiplayer battles—spanning WebGL rendering, Deno services, serverless Postgres, and automated browser testing.
BUILD/02 · FIELD GUIDE
The Geometry of Intelligence
Riemannian and information geometry for AI, made visual.
↗
BUILD/03 · INTERACTIVE ESSAY
Particles → Probability
From interacting particles to the mathematics of Fields Medalist Yu Deng.
↗
BUILD/04 · DEVELOPER TOOL
File → Prompt
A private-by-design browser tool that orders project files and turns them into one
structured prompt. Everything stays local.
↗
BUILD/05 · CONVERSATIONAL WEB
PengchengGPT
A conversational interface to my work, background, and public projects.
↗
inspect source: PengchengGPT repository ↗
side.quest / itch.io
offscreen.06 / human context
Music, speculative writing, browser experiments, and the story encoded in a name.
SELECTED PERSONAL MEDIA
A guitar recording away from the research loop. The embedded player loads only after you choose it, so Bilibili receives no request on the initial page load.
Watch on Bilibili ↗NAME / ORIGIN
In 鹏, 鸟 gives the bird meaning while 朋 supplies the sound. In the animation, the two 月 forms remain side by side within a large 朋 contour, roughly matching 鸟 in height, and become the feather framework of one huge open left wing. A recognizable bird silhouette forms the body while a lighter, semi-transparent 鸟 remains visible inside it; the character’s inner dot is highlighted as the eye. Together they form a visual metaphor for 鹏, the enormous bird of Chinese mythology, setting out on a ten-thousand-li journey. My parents chose Pengcheng as a hope for an ambitious path and a wide horizon. In English, I pronounce Xu like “Hsu.”
AUDIO / VIDEO
I sing, play guitar, and record occasional technology talks.
WORDS / SPECULATIVE
祂 and Father Sun are two forms of the same science-fiction project.
OFFLINE / INTERESTS
Basketball, tennis, table tennis, swimming, reading, science fiction, and learning how things work. Richard Feynman and Tsung-Dao Lee remain enduring inspirations.
I try to make work that contributes something positive and outlasts the moment in which it was made. The full personal note remains in the academic archive.
channel.07 / open
I welcome conversations about large language models, reinforcement learning, reliable AI evaluation, genomics, distillation, scientific machine learning, and unusually ambitious web experiments.
Choose 30 minutes for a focused question, or one hour for a deeper research, paper, or project conversation. The hour-long calendar also includes weekends.
Live booking is handled by Google Calendar; the email fallback goes to my UC Irvine inbox. The calendar loads only after you open this panel.
Open this panel to see current weekday and weekend availability.