Content Adversarial Red Team Analyst- English(Philippines)
WelocalizeJob Overview
We are seeking creative and analytical Content Adversarial Red Team Analysts to test AI platforms and models against content safety policies and compliance requirements.
In this role, you will design and execute challenging test scenarios to identify gaps in how AI systems respond to complex, unusual, or unexpected inputs. You will explore edge cases and different approaches to assess whether platforms and models consistently follow defined safety policies and compliance boundaries.
An ideal candidate is curious, analytical, and able to think from different user perspectives. You should have a strong understanding of content safety and policy and be comfortable exploring complex scenarios to identify weaknesses, inconsistencies, or gaps in system behaviour.
Your work will help identify areas of improvement and support the development of safer and more reliable AI systems.
Project Details
- Contract Type: Freelance, with the potential to convert to a full-time role.
- Pay Rate: US$10 per hour
- Location: Philippines
- Language: English
Responsibilities
- Design and execute authorized adversarial test scenarios to evaluate AI systems against content safety policies and compliance requirements.
- Develop diverse and challenging prompts, inputs, and scenarios to test system behaviour.
- Explore edge cases, unusual inputs, and complex situations that may reveal gaps in safety controls or policy enforcement.
- Evaluate AI responses and identify potential weaknesses, inconsistencies, or policy-related concerns.
- Test how systems respond to different forms of language, context, intent, and user behaviour.
- Identify recurring patterns or scenarios that may require further testing or improvement.
- Document test scenarios, system responses, findings, and supporting evidence clearly.
- Apply defined testing guidelines and project requirements consistently.
- Review complex cases and use sound judgment when assessing system behaviour.
- Maintain high levels of quality, accuracy, and attention to detail while completing assigned tasks.
Required Qualifications
- Educational background or equivalent experience in Trust & Safety, Content Safety, Policy, Linguistics, Communications, Journalism, Research, AI Evaluation, or a related field.
- Experience in content safety, trust and safety, AI evaluation, content moderation, policy enforcement, quality assurance, or related areas.
- Strong understanding of content safety policies, policy enforcement, and common safety risks.
- Strong understanding of written language, context, intent, and different ways users may communicate.
- Ability to think creatively and develop challenging, unusual, or unexpected test scenarios.
- Strong analytical and critical-thinking skills.
- Ability to identify patterns, weaknesses, inconsistencies, and gaps in system behaviour.
- Comfortable working with complex or ambiguous situations and making informed decisions based on defined requirements.
- Ability to understand and consistently apply detailed testing guidelines and policies.
- Strong written English comprehension and communication skills.
- Strong documentation skills and attention to detail.
- Familiarity with AI systems, large language models, adversarial testing, red teaming, or AI safety is preferred.