Jobs›Safety Specialist

AI Safety Specialist - Bilingual

mercor · Work from home
Work from home
PayPay not listed
WhereWork from home
TypeFull timeMid-level
Posted21 Sep15 days ago, via Himalayas
Skills they list12 named
AI SafetyRed TeamingAdversarial Machine LearningAI ResearchContent SafetyAI Safety SpecialistBilingual AI SpecialistAI Safety ExpertAI Safety PractitionerAI Safety AnalystMultilingual AI SpecialistAI Safety Researcher
About this job

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.

Position
AI Safety Experts — English & Malayalam
Type
Contract
Compensation
$16–$22/hour
Location
Remote

Role Responsibilities

Red team conversational AI models and agents to identify jailbreaks, prompt injections, and misuse cases.

Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.

Apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing.

Document reproducibly by producing reports, datasets, and attack cases that customers can act on.

Work independently and asynchronously to meet deadlines while improving AI model performance.

Qualifications

Must-Have

Fluent Language Skills Required:English & Malayalam. Native fluency in English and Malayalam is required.

Strong judgment about language and content accuracy.

Rigorous attention to detail and ability to notice subtle errors.

Structured approach to work following guidelines and quality standards.

Clear communication skills for technical and non-technical audiences.

Adaptability across projects, task types, and customers.

Preferred

Experience in Adversarial ML
jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
Cybersecurity skills
penetration testing, exploit development, reverse engineering.

Understanding of socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.

Creative probing skills
psychology, acting, writing for unconventional adversarial thinking.

Application Process (Takes 20–30 mins to complete)

Upload resume

AI interview based on your resume

Submit form

Resources & Support

For details about the interview process and platform information, please check

For any help or support, reach out to

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Originally posted on Himalayas

Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →

Apply on Himalayas
Opens himalayas.app in a new tab