Copy of LLM Model Response Evaluation
- Role
- AI / ML
- Experience
- Mid
- Employment
- Contract
Open to CA only. Set where you work from to check your eligibility.
No BS summary
LLM/GenAI data evaluation expert with 3+ years of hands-on experience and a Bachelor's Degree. Must be comfortable evaluating AI outputs across text, images, audio, video, HTML widgets, PDFs, and other modalities. Remote contract work with variable workload and no guaranteed hours.
Core skills
Required skills
- Contract
Our client, a global technology company that helps businesses build, train, and manage AI systems is looking for experts to evaluate model-generated content against defined quality rubrics such as factuality, consistency, aesthetics, and other evaluation criteria.
- Evaluating UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities.
- The work may involve text, images, audio, video, HTML widgets, PDFs, or combinations of these modalities.
- The work is domain-agnostic and may cover topics across arts, culture, history, science, engineering, and more.
- Resources will be expected to independently research unfamiliar topics using trusted sources before making evaluation decisions.
- Each task will include detailed project guidelines within the evaluation platform.
- 3+ years of hands-on experience in LLM / GenAI data evaluation.
- Bachelor's Degree required
- Ability to research unfamiliar topics using trusted sources and make well-supported judgments.
- Comfortable evaluating content across multiple modalities
Flexible and remote work Variable workload: Accept or decline tasks based on your availability No guaranteed hours: Workload may vary weekly
By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply
I'm interested
I'm interested
Privacy Notice
What you'll do
- Evaluate model-generated content against defined quality rubrics such as factuality, consistency, aesthetics, and other evaluation criteria.
- Evaluate UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities.
- Work with text, images, audio, video, HTML widgets, PDFs, or combinations of these modalities.
- Independently research unfamiliar topics using trusted sources before making evaluation decisions.
- Follow detailed project guidelines within the evaluation platform.
What they require
- 3+ years of hands-on experience in LLM / GenAI data evaluation.
- Bachelor's Degree required.
- Ability to research unfamiliar topics using trusted sources and make well-supported judgments.
- Comfortable evaluating content across multiple modalities.
- Variable workload: Accept or decline tasks based on your availability.
- No guaranteed hours: Workload may vary weekly.
Benefits
- Flexible and remote work
- Accept or decline tasks based on your availability
One of Upwork's largest clients, an American multinational technology company, is in search of an M365 Administrator.
What people say about this company
3.1/ 5