Preference Model · Member of Technical Staff - Research & Post-training

Measure whether an RL environment improved real capability.

Mohamed A M Elansary, PhD — multimodel evaluation under uncertainty, agent evaluation sets, and coverage-gap measurement for post-training loops.

Post-training evaluationAgent evalsUncertainty quantificationCoverage gaps

Evaluation under uncertainty

  • Six-plus years of multimodel, multi-basin forecast experiments across hydroclimates on Linux/HPC.
  • Compared statistical and physically based stacks, quantified uncertainty, and reported regime-dependent failure modes rather than a single flattering score.
  • That is the measurement analogue of validating data quality, surfacing task-coverage gaps, and closing a feedback loop between environment design and capability.

Agent evaluation sets

  • Production GPT, Claude, and Gemini agent workflows with retrieval, routing, tenant isolation, provenance, and regression evaluation sets at Vertexium, including a multi-tenant conversational receptionist.
  • That maps to inspecting whether an agent trajectory reflects intended behavior. It is not reward-model training, RLHF, or Preference Model environment ownership.
  • Python scientific/HPC pipelines on UCAR and TACC Stampede; not a claim of PyTorch or JAX LLM post-training.

Proposed first contribution

For one proprietary RL environment already in use, define what capability the environment is supposed to teach versus a score that is easy to move. Write a small failure taxonomy: metric movement without a capability change, uncovered task slices, an evaluation signal that disagrees with intended behavior, a data-quality failure that looks like a training win. Stand up a small evaluation set with provenance, compare simple baselines, attach uncertainty, and write a clear report of coverage gaps before expanding Verl or OpenRLHF training-infrastructure work. This is a proposed measurement approach, not a claim of prior Preference Model-internal work, RLHF, preference-model training, or invented metrics.

Honest fit boundary

RLHF and preference-model training depth is a stretch. I have not run end-to-end LLM post-training of 7B-or-larger models, trained reward or preference models, used Verl or OpenRLHF, designed Preference Model-internal RL environments, or claimed Preference Model-internal work, and I do not invent metrics or safety research. The credible contribution is evaluation under uncertainty, agent evaluation harnesses, scientific/HPC rigor, and production data pipelines.

Role and location

Member of Technical Staff - Research & Post-training · San Francisco · OnSite. Ashby lists workplaceType OnSite and isRemote false. JD quote: “Visa sponsorship & relocation support available.” Willing to relocate to San Francisco. Remote work is not asserted.

Posting compensation: “$200K – $350K • Offers Equity”. Clearance / citizenship / ITAR: none found in the live Ashby posting (re-verified 2026-09-03). · Official role posting