Skip to main content
Phonely

Applied Machine Learning Researcher

RemoteAustralia only· Prefers Australia
Published
Role
AI / ML
Employment
Full-time
Salary not disclosed
Check eligibility

Open to AU only · Prefers Australia. Set where you work from to check your eligibility.

The listing prefers candidates in Australia.

No BS summary

Applied ML researcher for conversational voice AI, focused on fine-tuning and evaluating language models in production. Must have strong ML/LLM experience, Python, real-world datasets, evaluation systems, and clear communication. Remote role, preferably based in Australia.

Core skills

PythonLLM fine-tuningLLM evaluation

Optional skills

SFTDPORFTGRPORLHFLoRAQLoRAvLLM

What you'll do

  • Research and improve LLM behavior for real-time voice conversations
  • Design and run fine-tuning experiments across data, model, and evaluation strategies
  • Build evaluation frameworks for model quality, workflow-following, naturalness, reliability, task completion, and overall conversation quality
  • Analyze production conversations to find failure modes and opportunities to improve
  • Develop data curation, labeling, and synthetic data strategies
  • Compare model architectures, training approaches, prompts, and datasets
  • Investigate regressions and explain clearly why model behavior improves or degrades
  • Work with engineering to deploy research improvements safely and efficiently
  • Help define model release criteria, eval gates, and quality benchmarks
  • Partner closely with engineering, product, QA, and customer-facing teams to understand where models break in the real world
  • Improve training data and ship better models
  • Write Python and analyze real production conversations

What they require

  • Strong experience with machine learning and modern language models
  • Hands-on experience fine-tuning, evaluating, or adapting language models
  • Strong Python skills and comfort working with messy, real-world datasets
  • Ability to design rigorous experiments and interpret the results clearly
  • Experience building or improving evaluation systems for AI models
  • Strong analytical skills and the ability to debug model behavior
  • Clear written and spoken communication, so you can explain findings to technical and non-technical teammates alike
  • A practical mindset: you care about production impact, latency, reliability, and customer outcomes
  • This is a hands-on research and engineering role
  • Preferred: A PhD in machine learning, NLP, or a related field
  • Preferred: Experience with conversational AI, voice AI, or customer-support automation
  • Preferred: Familiarity with model serving, inference optimization
  • Preferred: Experience with synthetic data generation and data quality pipelines
  • Preferred: Experience working in a startup or fast-moving product environment

Benefits

  • Remote team
  • A few times a year the team gets together in person and rents out Airbnbs in places such as the Rocky Mountains, Costa Rica, and Indonesia

Phonely builds conversational voice AI agents for high volume phone workflows, used to qualify leads, book appointments, route calls, and resolve customer conversations in production.

AIStartup

Details

Apply routeDom
Salary not disclosed