Skip to main content
Salvo Software

AI Developer

RemoteUnited States only
Published
Role
AI / ML
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Hands-on AI/ML backend developer for LLM training, fine-tuning, RAG, MCP, and offline/on-prem deployments. Needs strong Python, ML frameworks, vector databases, Docker/Git, CUDA/GPU optimization, and experience with air-gapped environments. US-based role.

Core skills

RAGMCPLLM fine-tuning

Required skills

PythonPyTorch/TensorFlowscikit-learnpandasPostgreSQL/MySQLDockerGitHugging Face TransformersHugging Face DatasetsXMLXSDvLLM/TGI/OllamaGGUF/GPTQ/AWQCUDAHyDE

Optional skills

ML model registriesAWS

What you'll do

  • Train and fine-tune LLMs using supervised fine-tuning (SFT).
  • Work with open-source models such as LLaMA, Mistral, Qwen, and similar architectures.
  • Build LoRA / Q-LoRA pipelines for efficient fine-tuning.
  • Implement and optimize data preprocessing workflows, including tokenization and long-context handling.
  • Use and extend Hugging Face Transformers & Datasets for training and inference.
  • Parse and process structured and semi-structured data, including XML/XSD files.
  • Implement document parsing solutions for Office formats (python-docx, OpenXML).
  • Design and implement end-to-end Retrieval-Augmented Generation (RAG) pipelines for document-grounded question answering and knowledge retrieval.
  • Build and maintain vector stores and embedding pipelines using tools such as FAISS, Chroma, Weaviate, or pgvector.
  • Optimize retrieval strategies including hybrid search, re-ranking, and chunking approaches tailored for domain-specific corpora.
  • Develop and maintain MCP (Model Context Protocol) server integrations to enable LLMs to interact dynamically with tools, APIs, and external data sources.
  • Design agentic workflows that leverage MCP to give models structured access to internal systems and context in a controlled, auditable manner.
  • Deploy, run, and maintain models fully offline and in air-gapped environments.
  • Perform model optimization and quantization (GGUF, GPTQ, AWQ, bitsandbytes).
  • Build and maintain inference systems using frameworks like vLLM, TGI, and Ollama.
  • Optimize GPU usage (CUDA, cuDNN, VRAM-aware batching).
  • Maintain local CI/CD pipelines for ML models without cloud dependencies.
  • Manage local model registries, versioning, and artifacts.
  • Ensure RAG and MCP components are fully operational in offline and restricted network environments.
  • Build backend services in Python for ML training and inference workflows.
  • Work with relational databases (Postgres/MySQL) and vector databases for RAG storage layers.
  • Use Docker and Git for reliable development and deployment pipelines.
  • Use Azure DevOps for CI/CD, including local runners when applicable.

What they require

  • Strong experience in Python for backend and machine learning development.
  • Expertise with ML frameworks such as PyTorch or TensorFlow, along with scikit-learn and pandas.
  • Solid knowledge of Postgres or MySQL for data storage.
  • Experience with Docker and Git.
  • Hands-on experience with LLM training, fine-tuning, and optimization.
  • Experience with Hugging Face Transformers & Datasets.
  • Familiarity with XML/XSD and Office document parsing tools.
  • Experience deploying models with vLLM, TGI, or Ollama.
  • Understanding of quantization techniques such as GGUF, GPTQ, or AWQ.
  • Experience with GPU optimization and the CUDA stack.
  • Experience building solutions for offline, on-prem, and air-gapped environments.
  • Hands-on experience designing and implementing RAG pipelines, including embedding models, vector stores, and retrieval optimization strategies.
  • Experience building or integrating MCP (Model Context Protocol) servers to connect LLMs with external tools, APIs, and structured data sources.
  • Experience with advanced RAG techniques such as HyDE or multi-hop retrieval.
  • Experience building agentic systems using MCP in production or near-production environments.
  • Preferred: Experience with secure environments, restricted networks, or enterprise compliance requirements.
  • Experience discussing complex technical topics with both technical and non-technical stakeholders.

Salvo Software is a tight-knit team of Software builders, problem-solvers, and technology enthusiasts on a mission to help businesses grow through custom software. Headquartered in Vancouver, WA, with engineering talent across the globe, we deliver enterprise-grade solutions with the speed and care that only a passionate, people-first team can offer.

🇨🇦 CanadaSoftware DevelopmentStartup
Salary not disclosed