
AI&
Member of Technical Staff - Post Training
Japan (Hybrid)RemotePosted 27 days ago¥9,000,000 – ¥9,000,000
Full TimeSeniorRemoteJP
See how this job matches your profile
Sign in for an AI-powered fit score, breakdown, and a tailored resume.
Job Description
This is both a research and an engineering role. You will own post-training end to end for ai&’s internal custom models and for enterprise customers who need models adapted to their domains. That mean
Key Highlights
- Post-Training Pipeline Engineering Build and maintain the full post-training infrastructure including SFT, preference alignment, reward model pipelines, experiment tracking, and evaluation infrastructure. Own this stack for both internal model development and enterprise engagements.
- Enterprise Post-Training Ownership Act as the technical owner for enterprise customer post-training engagements. Translate customer requirements into concrete post-training specifications, run the workflows, design task-specific evaluations, and feed learnings back into core pipelines.
- Data Generation & Quality Design and build synthetic data pipelines that support post-training and RL at scale. Own generation, filtering, and quality assessment workflows. Strong intuition for data quality is non-negotiable.
- Continual Learning Develop the methodologies and infrastructure that allow ai& models to keep improving over time without catastrophic forgetting. Design training regimes and evaluation protocols that support ongoing model development as new data and feedback accumulates.
- Evaluation & Benchmarking Design task-specific evaluations that go beyond standard benchmarks. Interpret results honestly, catch regressions before they reach production, and use findings to drive concrete improvements.
Qualifications
Required Qualifications
- Reinforcement Learning in Practice You have actually run RL on language models. You have implemented reward models, dealt with reward hacking, tuned KL penalties, and shipped models that are meaningfully better as a result. You understand the theory and you have applied it.
- Post-Training Engineering Depth Hands-on experience with data generation and evaluation for LLM post-training. You have run SFT, preference alignment, and RL workflows on real models and you know where these pipelines break.
- Framework Proficiency Strong Python and PyTorch proficiency with hands-on experience optimizing training pipelines. Experience with DeepSpeed, FSDP, vLLM, or similar frameworks for efficient model training and inference.
- Data Quality Instinct Strong intuition for what good training data looks like. Experience designing and executing data generation, filtering, curation, and quality assessment processes at scale.
- End-to-End Thinking You reason across data generation, training, alignment, and evaluation as a single system. You do not optimize one stage in isolation from the others.
- Customer and Communication Fluency Comfortable working directly with enterprise customers. You can translate between customer needs and internal technical teams, push back when needed, and be trusted as the technical owner of a delivery.
- Continual Learning Familiarity Familiar with the challenges of continual and lifelong learning in neural networks. You have thought seriously about catastrophic forgetting and how to build models that stay current without degrading.
- Great Team Spirit A mission-driven approach to engineering, valuing clear communication, hands-on execution, and collective success over individual silos.
Skills & Technologies
GoPythonPyTorch
About the Company
AI&
View company profile →
Interested in this role?
Sign in or create a free account to see how this job matches your skills, apply with one click, and let our AI tailor your resume.
Sign in to applyAI-powered resume optimization
Save and track your applications
Job Details
Employment Type
Full Time
Experience Level
Senior
Salary Range
¥9,000,000 – ¥9,000,000
Location
Japan (Hybrid)
Work Mode
Remote
Posted
27 days ago
Country
JP