LLM Model Response Evaluation
- Role
- AI / ML
- Experience
- Mid
- Employment
- Contract
Open to US only. Set where you work from to check your eligibility.
No BS summary
Contract role for someone with 3+ years of hands-on LLM/GenAI data evaluation experience and a bachelor's degree. Must be comfortable evaluating model-generated content across text, images, audio, video, HTML widgets, PDFs, and other modalities.
- Contract
Our client, a global technology company that helps businesses build, train, and manage AI systems is looking for experts to evaluate model-generated content against defined quality rubrics such as factuality, consistency, aesthetics, and other evaluation criteria.
- Evaluating UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities.
- The work may involve text, images, audio, video, HTML widgets, PDFs, or combinations of these modalities.
- The work is domain-agnostic and may cover topics across arts, culture, history, science, engineering, and more.
- Resources will be expected to independently research unfamiliar topics using trusted sources before making evaluation decisions.
- Each task will include detailed project guidelines within the evaluation platform.
- 3+ years of hands-on experience in LLM / GenAI data evaluation.
- Bachelor's Degree required
- Ability to research unfamiliar topics using trusted sources and make well-supported judgments.
- Comfortable evaluating content across multiple modalities
Flexible and remote work Variable workload: Accept or decline tasks based on your availability No guaranteed hours: Workload may vary weekly
By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply
I'm interested
I'm interested
Privacy Notice
What you'll do
- Evaluate model-generated content against defined quality rubrics such as factuality, consistency, aesthetics, and other evaluation criteria.
- Evaluate UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities.
- Work with text, images, audio, video, HTML widgets, PDFs, or combinations of these modalities.
- Independently research unfamiliar topics using trusted sources before making evaluation decisions.
- Follow detailed project guidelines within the evaluation platform.
What they require
- 3+ years of hands-on experience in LLM / GenAI data evaluation.
- Bachelor's Degree required.
- Ability to research unfamiliar topics using trusted sources and make well-supported judgments.
- Comfortable evaluating content across multiple modalities.
- Variable workload: Accept or decline tasks based on your availability.
- No guaranteed hours: Workload may vary weekly.
Benefits
- Flexible and remote work
One of Upwork's largest clients, an American multinational technology company, is in search of an M365 Administrator.
What people say about this company
3.1/ 5