AI Research Engineer (Pre-training - LLM & Multi-Modal)
- Role
- AI / ML
The listing doesn't say where it hires from. Check the description or the employer's site before applying.
No BS summary
AI Research Engineer with deep expertise in LLM and Multi-Modal architectures, and a strong grasp of pre-training optimization. Must have hands-on, research-driven experience with large-scale LLM or Multi-Modal pre-training runs on distributed servers with thousands of GPUs. PhD in NLP, Machine Learning, or related field preferred.
Core skills
Required skills
Optional skills
Required languages
Join Tether and Shape the Future of Digital Finance At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction. Innovate with Tether Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT , relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services. But that’s just the beginning: Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities. Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET , our flagship app that redefines secure and private data sharing. Tether Education : Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity. Tether Evolution : At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways. Why Join Us? Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry. If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you. Are you ready to be part of the future? About the job As a member of the AI model team, you will drive innovation in architecture development for cutting-edge models of various scales, including small, large, and multi-modal systems. Your work will enhance intelligence, improve efficiency, and introduce new capabilities to advance the field. You will have a deep expertise in Large Language Model (LLM) and Multi-Modal architectures, a strong grasp of pre-training optimization, and a hands-on, research-driven approach. Your mission is to explore and implement novel techniques and algorithms that lead to groundbreaking advancements: multi-modal data curation and alignment, strengthening baselines, and identifying and resolving existing pre-training bottlenecks to push the limits of cross-modal AI performance. Responsibilities Large-Scale Pre-Training: Conduct foundational pre-training for LLMs and Multi-Modal models (integrating text, vision, audio, or other modalities) on large, distributed servers equipped with multi-nodes & thousands of NVIDIA GPUs. Architecture & Alignment Innovation: Design, prototype, and scale innovative architectures, tokenizers, and cross-modal alignment layers to enhance model intelligence and multi-modal understanding. Data Strategy: Source, filter, and curate massive-scale textual and multi-modal datasets, establishing robust data pipelines for efficient pre-training. Experimental Research: Independently and collaboratively execute experiments, analyze results, and refine training methodologies for optimal performance and token efficiency. Optimization & Debugging: Investigate, debug, and eliminate bottlenecks in model efficiency, computational performance, and multi-modal alignment stability during long training runs. System Scalability: Contribute to the advancement of distributed training systems to ensure seamless scalability and hardware efficiency on target platforms. A degree in Computer Science or related field. Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences). Hands-on experience contributing to large-scale LLM or Multi-Modal pre-training runs on large, distributed servers equipped with thousands of NVIDIA GPUs, ensuring scalability and impactful advancements in model performance. Familiarity and practical experience with large-scale, distributed training frameworks, libraries and tools. Deep knowledge of state-of-the-art transformer and non-transformer modifications aimed at enhancing intelligence, efficiency and scalability. Strong expertise in PyTorch and Hugging Face libraries with practical experience in model development, continual pretraining, and deployment. Important information for candidates Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles: Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website. Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms. Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately. When in doubt, feel free to reach out through our official website.
What you'll do
- Conduct foundational pre-training for LLMs and Multi-Modal models (integrating text, vision, audio, or other modalities) on large, distributed servers equipped with multi-nodes & thousands of NVIDIA GPUs.
- Design, prototype, and scale innovative architectures, tokenizers, and cross-modal alignment layers to enhance model intelligence and multi-modal understanding.
- Source, filter, and curate massive-scale textual and multi-modal datasets, establishing robust data pipelines for efficient pre-training.
- Independently and collaboratively execute experiments, analyze results, and refine training methodologies for optimal performance and token efficiency.
- Investigate, debug, and eliminate bottlenecks in model efficiency, computational performance, and multi-modal alignment stability during long training runs.
- Contribute to the advancement of distributed training systems to ensure seamless scalability and hardware efficiency on target platforms.
- Apply low-bit quantization to reduce model size and inference latency for generative AI models (LLMs, VLMs, multimodal) while maintaining accuracy and output quality.
- Leverage knowledge distillation to transfer capabilities from larger teacher models to smaller student models, enabling efficient multimodal reasoning across text, image, and audio inputs.
- Implement pruning techniques to remove redundant parameters and attention heads, reducing computational overhead without sacrificing task performance.
- Analyze trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning methods; propose improvements based on empirical findings.
- Research and apply mixed-precision quantization and other advanced compression strategies (e.g., adaptive pruning schedules, distillation with intermediate feature matching) to optimize the accuracy–performance balance.
- Stay current with the latest research in model compression, including emerging techniques for multimodal and generative architectures.
- Document methodologies, experiments, and results clearly to support reproducibility, internal collaboration, and stakeholder communication.
- Author technical papers and publish findings in top-tier conferences (e.g., NeurIPS, ICML, ICLR, CVPR, ACL, AAAI) to advance the field of model compression for multimodal AI.
- Design and deploy state-of-the-art model serving architectures that deliver high throughput and low latency while optimizing memory usage.
- Ensure these pipelines run efficiently across diverse environments, including resource-constrained devices and edge platforms.
- Establish clear performance targets such as reduced latency, improved token response, and minimized memory footprint.
- Build, run, and monitor controlled inference tests in both simulated and live production environments.
- Track key performance indicators such as response latency, throughput, memory consumption, and error rates, with special attention to metrics specific to resource-constrained devices.
- Document iterative results and compare outcomes against established benchmarks to validate performance across platforms.
- Identify and prepare high-quality test datasets and simulation scenarios tailored to real-world deployment challenges, specifically those encountered on low-resource devices.
- Set measurable criteria to ensure that these resources effectively evaluate model performance, latency, and memory utilization under various operational conditions.
- Analyze computational efficiency and diagnose bottlenecks in the serving pipeline by monitoring both processing and memory metrics.
- Address issues such as suboptimal batch processing, network delays, and high memory usage to optimize the serving infrastructure for scalability and reliability on resource-constrained systems.
- Work closely with cross-functional teams to integrate optimized serving and inference frameworks into production pipelines designed for edge and on-device applications.
- Define clear success metrics such as improved real-world performance, low error rates, robust scalability, optimal memory usage and ensure continuous monitoring and iterative refinements for sustained improvements.
- Conduct research on reinforcement learning algorithms for multimodal models, including diffusion-based approaches for image autoregressive models for multimodal understanding, and unified frameworks that integrate multiple modalities.
- Design and build reinforcement learning infrastructure that supports scalable, distributed training across multimodal systems while maintaining efficiency and reliability.
- Develop and refine reward modeling strategies that improve training stability, align model behavior with desired outcomes, and mitigate reward hacking and related failure modes.
- Create and curate multimodal simulation environments and datasets to support robust training, evaluation, and benchmarking of reinforcement learning systems.
- Design and conduct rigorous benchmarking and evaluation protocols to measure model performance, track progress against baselines, and validate improvements across multimodal tasks.
- Analyze and optimize policy performance across modalities by identifying bottlenecks in training, credit assignment, and cross-modal alignment.
- Investigate and develop next-generation reinforcement learning paradigms that more effectively learn from environment feedback, with the goal of achieving superior state-of-the-art (SOTA) performance.
- Publish research findings in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV etc.
- Conduct end-to-end research and engineering on vision-language models, covering training, evaluation, and optimization across the full model development lifecycle.
- Design and implement post-training pipelines including supervised fine-tuning, knowledge distillation, and reinforcement learning from human feedback.
- Develop and maintain high-quality multimodal datasets, including data curation, filtering, and balancing for domain-specific tasks.
- Drive model efficiency and deployability, adapting models for resource-constrained environments using compression and optimization techniques.
- Design and implement evaluation frameworks and benchmarks to measure model performance, robustness, and real-world task success.
- Build and scale training workflows across distributed GPU infrastructure.
- Identify and resolve bottlenecks in training pipelines to achieve state-of-the-art model quality on target benchmarks.
- Contribute to and leverage open-source ecosystems including models, datasets, and tooling to accelerate development.
- Stay current with the latest research in multimodal learning and vision-language systems, translating relevant findings into practical improvements.
- Publish research findings in top-tier AI conferences and journals where applicable.
What they require
- A degree in Computer Science or related field.
- Ideally PhD in NLP, Machine Learning, or a related field, complemented by a solid track record in AI R&D (with good publications in A* conferences).
- Hands-on experience contributing to large-scale LLM or Multi-Modal pre-training runs on large, distributed servers equipped with thousands of NVIDIA GPUs, ensuring scalability and impactful advancements in model performance.
- Familiarity and practical experience with large-scale, distributed training frameworks, libraries and tools.
- Deep knowledge of state-of-the-art transformer and non-transformer modifications aimed at enhancing intelligence, efficiency and scalability.
- Strong expertise in PyTorch and Hugging Face libraries with practical experience in model development, continual pretraining, and deployment.
- Excellent English communication skills
- Experience with PyTorch deep learning frameworks or equivalent frameworks
- Hands-on experience with model quantization including both Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ).
- Research and hands-on experience with knowledge distillation for compressing large models into smaller, efficient ones.
- Research and hands-on experience with model pruning for compressing large models into smaller, efficient ones.
- Solid understanding of neural network architectures and training processes – Including transformers (e.g., LLMs, VLMs), backpropagation, optimization, and fine-tuning techniques.
- Familiarity with C++ is a plus (especially for implementing low-level quantization kernels or inference optimizations).
- Must have knowledge of Metal Shading Language (MSL).
- You should be comfortable writing custom compute shaders from scratch.
- Proven experience in low-level kernel optimizations and inference optimization on mobile devices is essential.
- Your contributions should have led to measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications, particularly on resource-constrained devices and edge platforms.
- A deep understanding of modern model serving architectures and inference optimization techniques is required.
- This includes state-of-the-art methods for achieving low-latency, high-throughput performance, and efficient memory management in diverse, resource-constrained deployment scenarios.
- Must have strong expertise in writing GPU kernels for mobile devices (i.e., smartphones) as well as a deep understanding of model serving frameworks and engines.
- Practical experience in developing and deploying end-to-end inference pipelines, from optimizing models for efficient serving to integrating these solutions on resource-constrained devices is required.
- Demonstrated ability to apply empirical research to overcome challenges in model serving, such as latency optimization, computational bottlenecks, and memory constraints.
- You should be proficient in designing robust evaluation frameworks and iterating on optimization strategies to continuously push the boundaries of inference performance and system efficiency.
- Distributed Inference Systems: Designing and optimizing high-performance inference engines using techniques like Tensor Parallelism, Pipeline Parallelism, and Expert Parallelism to handle massive models on GPU clusters.
- Deep understanding of the math and structure behind Diffusion Models and Vision Transformers
- Understanding of Pruning, Quantization, Flash attention, KV Cache, Speculative Decoding (Eagle) etc.
- Designing and optimizing high-performance inference engines using techniques like Tensor Parallelism, Pipeline Parallelism, and Expert Parallelism to handle massive models on GPU clusters.
- A Master's degree in Computer Science or a related field is required; a PhD in Machine Learning, NLP, Computer Vision, or a closely related discipline is preferred, along with a strong track record of AI research and publications in top-tier conferences.
- Proven experience running large-scale reinforcement learning experiments in multimodal and vision-centric systems, including online RL settings, with demonstrated impact on domain-specific decision-making and measurable improvements in policy performance.
- Deep understanding of reinforcement learning algorithms and optimization methods applied to vision and multimodal learning problems, with a focus on improving policy stability, exploration, and sample efficiency in complex, high-dimensional environments involving images, video, and other modalities.
- Strong proficiency in PyTorch and deep learning frameworks for vision and multimodal AI, with hands-on experience building end-to-end RL pipelines covering simulation, training, evaluation, and deployment in production-grade systems.
- Demonstrated ability to apply empirical research to solve core RL challenges in multimodal and vision tasks, such as sample inefficiency, exploration-exploitation tradeoffs, and training instability, along with experience designing robust evaluation frameworks and iterating on algorithmic improvements to advance agent performance.
- Proven track record of research publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV etc.
- Degree in Computer Science, Machine Learning, or a related field; MS/PhD preferred.
- Strong experience with multimodal post-training workflows including supervised fine-tuning, knowledge distillation, and reinforcement learning from feedback.
- Hands-on experience with parameter-efficient fine-tuning and distributed training frameworks.
- Demonstrated ability to build and improve vision-language models with measurable results on standard benchmarks or real-world tasks.
- Experience adapting models for resource-constrained environments.
- Proven open-source contributions in multimodal AI on GitHub or HuggingFace.
- Publications at top AI conferences (NeurIPS, ICML, ICLR, CVPR, ECCV etc.)
- A deep understanding of modern model serving architectures and inference optimization techniques is required. This includes state-of-the-art methods for achieving low-latency, high-throughput performance, and efficient memory management in diverse, resource-constrained deployment scenarios.
Benefits
- Our team is a global talent powerhouse, working remotely from every corner of the world.
- If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards.
- We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
- If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Tether builds digital finance products including USDT, digital asset tokenization services, energy solutions for Bitcoin mining, AI and peer-to-peer technology products, and digital education initiatives.