AI Model Policy Trainer, Image Evaluation - Seattle Onsite
Handshake Seattle, WAEst. Est. USD 95,000–130,000 / yearMid
Estimated range based on role, country and industry — not published by the company.
Key requirements
- Workday
- Audit
About Handshake
Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.
Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise.
Role Details
Location: Onsite in Seattle, WA, Monday-Friday.
Compensation: $36-$72/hr. Placement within the range depends on experience.
Employer: TCWGlobal. This is a W-2 assignment supporting Handshake AI.
Employment: Full time, 40 hours per week, non-exempt and eligible for overtime pay.
Schedule: Monday-Friday, 8 a.m.-5 p.m. PT.
About the Role
As an AI Image Evaluator, you will help image generation models learn two things at once: what a good image is, and what an acceptable image is.
You will look at prompts and the images a model produced from them, then answer questions like: Did the image do what the prompt asked? Is it well made, or does it have the hands, lighting, text, and anatomy problems that give generated images away? Which of two images is better, and why? Does the image violate the customer's content policy, and if so, which category and how severely? Does it depict a real person, a protected brand, or a minor in a way the policy does not allow?
The interesting cases are the close ones. Two images that look nearly identical until you notice one has a logo in the background. A stylized nude that is fine as figure study and not fine with one change of pose. A prompt that asked for "a realistic photo of a senator" and a model that complied. A beautiful image that ignored half the prompt, next to an ugly one that nailed it.
We are looking for people who already see images critically, whether that came from photography, illustration, design, years inside Midjourney and Stable Diffusion, or moderating visual content at scale. You do not need all of these. You need one deep, and the judgment to learn the rest.
This is not rote annotation. Rubrics cannot anticipate every image, and good evaluators do not apply them mechanically. You will balance the rubric's text and intent with customer expectations, precedent, and team calibration, and you will explain your reasoning clearly enough that it can train a model.
What You Will Do
Evaluate generated images against their prompts for adherence, composition, realism, style consistency, and technical defects; maintain accuracy and consistency across repeated evaluations
Compare images side by side and select the stronger one with a clear, evidence-based rationale
Classify images against customer content policies covering sexual content, violence, hate symbols, real-person likeness, intellectual property, and depictions of minors
Select the most defensible classification when an image is genuinely ambiguous, and write concise rationales that cite rubric language and specific visual details
Distinguish "I do not like this" from "this fails the prompt" from "this violates policy," and keep those judgments separate while applying customer policy consistently
Write and refine prompts that probe where a model's quality or safety behavior breaks down
Identify rubric gaps, contradictions, and emerging edge cases, and raise them with project leads and policy teams
Participate actively in calibration discussions; challenge interpretations respectfully and update your judgment when stronger reasoning emerges
You May Be a Fit If
You have a trained eye from photography, illustration, concept art, art direction, retouching, photo editing, VFX, or visual design, and you can say precisely why one image is better than another
You use generative image tools heavily (Midjourney, Stable Diffusion, ComfyUI, Flux, DALL-E, Ideogram) and know their failure modes, their prompt quirks, and how their safety filters get bypassed
You have moderated or reviewed
See your match score for this role.
Xecodai maps the interview stages and shows what is preventing a 95% match.
