AI Data Engineer III
- Роль
- Инженерия данных
- Опыт
- Синьор
- Занятость
- Полная занятость
- Размер компании
- Крупная
Доступно для: BR only. Укажите, откуда вы работаете, чтобы проверить доступность.
Коротко по делу
Remote Brazil role for an experienced data engineer (5+ years) with 1–2+ years building AI/ML data pipelines and RAG systems. Hands‑on with Python, PostgreSQL (pgvector) and vector search/embeddings; experience integrating data from Salesforce/ServiceNow and building ETL/ELT pipelines. Preferred familiarity with Snowflake and enterprise data platforms.
Ключевые навыки
Обязательные навыки
Желательные навыки
Желательные языки
About Rimini Street, Inc. Rimini Street, Inc. (Nasdaq: RMNI), a Russell 2000® Company, is a global provider of end-to-end enterprise software support, products and services, the leading third-party support provider for Oracle and SAP software and a Salesforce and AWS partner. The Company has operations globally and offers a comprehensive family of unified solutions to run, manage, support, customize, configure, connect, protect, monitor, and optimize enterprise application, database, and technology software. To date, over 5,300 Fortune 500, Fortune Global 100, midmarket, public sector, and other organizations from a broad range of industries have relied on Rimini Street as their trusted enterprise software solutions provider. We are actively seeking an AI Data Engineer III. This is a remote role based in Brazil. Position Summary The AI Data Engineer is responsible for building the knowledge layer of Rimini Street’s Agentic ERP Platform—the data pipelines, RAG (Retrieval-Augmented Generation) systems, and embedding infrastructure that give AI agents access to the right information at the right time. This role owns how knowledge is ingested, processed, indexed, and retrieved to support intelligent agent behavior. Reporting to the Sr. Director, Engineering, this engineer designs the data architecture that powers agent intelligence—from extracting knowledge from Rimini Street’s 15+ years of support case history to building real-time retrieval systems for customer-specific context. The ideal candidate combines strong data engineering fundamentals with modern AI/ML knowledge, particularly in embeddings, vector search, and retrieval optimization. Essential Duties & Responsibilities RAG Pipeline Development Design and build RAG pipelines that retrieve relevant context from knowledge bases to augment AI agent responses. Implement chunking strategies optimized for different content types: support tickets, documentation, policies, transaction records, and email threads. Develop hybrid retrieval approaches combining dense embeddings, sparse search (BM25), and metadata filtering. Build query understanding and reformulation logic to improve retrieval relevance. Implement retrieval evaluation frameworks to measure and optimize precision, recall, and relevance. Design reranking pipelines that prioritize the most relevant results for agent consumption. Embedding & Vector Infrastructure Implement and manage vector storage using PostgreSQL with pgvector extension, including index optimization for search performance. Evaluate and select embedding models appropriate for enterprise content (technical documentation, business processes, ERP terminology). Build embedding pipelines that process documents at scale with appropriate batching and error handling. Implement incremental indexing strategies for real-time updates without full reprocessing. Design multi-tenant vector architectures that isolate customer data while enabling efficient search. Monitor and optimize vector search performance: latency, accuracy, and resource utilization. Data Ingestion & Processing Build data pipelines to ingest knowledge from diverse sources: Salesforce support tickets, ServiceNow cases, documentation repositories, email archives, and ERP transaction logs. Implement ETL processes that clean, normalize, and enrich raw data for AI consumption. Develop document processing pipelines: PDF extraction, HTML parsing, structured data normalization. Build connectors to source systems including Salesforce, ServiceNow, SharePoint, and Confluence. Implement data quality monitoring and alerting for ingestion pipelines. Design data lineage tracking to understand how knowledge flows from source to agent consumption. Knowledge Architecture Design the knowledge architecture that organizes information across the Four-Spoke model: Policy Intelligence, Institutional Memory, Rimini Collective Intelligence, and Intelligent Escalation. Build knowledge graphs and relationship models that capture connections between ERP concepts, processes, and solutions. Implement metadata taxonomies that enable filtered retrieval by ERP system, module, version, customer, and topic. Design versioning strategies for knowledge that evolves over time (policies, procedures, best practices). Build feedback loops that capture which retrieved content was useful vs. ignored, enabling continuous improvement. Cloud Data Platform Integration Integrate with Snowflake for large-scale data processing, leveraging Cortex AI capabilities where applicable. Build data pipelines that move and transform data between operational systems, Snowflake, and vector stores. Implement data access patterns that respect customer data isolation and security boundaries. Design efficient data synchronization between cloud data warehouse and real-time retrieval systems. Optimize query patterns for cost-effective data processing at scale. Experience 5+ years of data engineering experience, with at least 1-2 years focused on AI/ML data pipelines or RAG systems. Hands-on experience building and optimizing RAG pipelines in production environments. Strong experience with vector databases, embeddings, and similarity search. Experience with ETL/ELT pipelines and data integration from diverse source systems. Production experience with PostgreSQL and SQL-based data processing. Background in Python for data processing and pipeline development. Experience with enterprise data platforms (Snowflake, Databricks, or similar) preferred. Exposure to enterprise software, ERP systems, or support/ticketing systems preferred. Technical Skills Required Python for data engineering: pandas, data processing pipelines, async programming. PostgreSQL with strong SQL skills; experience with advanced features (JSONB, full-text search, extensions). Vector databases and embeddings: pgvector, or experience with Pinecone, Qdrant, Weaviate, or similar. RAG concepts: chunking strategies, embedding models, retrieval methods, reranking. ETL/ELT patterns and data pipeline orchestration (Airflow, Dagster, Prefect, or similar). Data modeling for both relational and document-oriented use cases. Git version control and CI/CD practices for data pipelines. Understanding of API integration for data extraction (REST, GraphQL). Preferred Experience with Snowflake, including Cortex AI features for vector search and ML functions. Familiarity with embedding models: OpenAI embeddings, Cohere, or open-source models (BGE, E5). Experience with LlamaIndex, LangChain, or Haystack for RAG pipeline development. Knowledge of document processing: PDF extraction (PyMuPDF, pdfplumber), HTML parsing, OCR. Experience with Salesforce and/or ServiceNow data extraction and APIs. Understanding of knowledge graphs and graph databases (Neo4j, or property graphs in PostgreSQL). Experience with data quality frameworks and monitoring tools. Familiarity with dbt for data transformation. Exposure to enterprise search systems (Elasticsearch, OpenSearch). Understanding of LLM fine-tuning and training data preparation. Skills & Competencies Strong analytical mindset with ability to understand complex data relationships and design efficient retrieval strategies. Data quality obsession; understands that agent intelligence is only as good as the underlying data. Systems thinker who designs for scale, reliability, and maintainability. Collaborative; works effectively with GenAI Engineers to understand retrieval requirements and optimize for agent consumption. Problem solver who can diagnose and resolve data pipeline issues quickly. Clear communicator; able to explain data architecture decisions to technical and non-technical stakeholders. Self-motivated and effective in a remote environment. Fluent in English (written and verbal). Desired Qualifications Bachelor’s or Master’s degree in Computer Science, Data Science, or related field. Experience in enterprise software companies or B2B SaaS platforms. Background in information retrieval, search systems, or NLP. Certifications in Snowflake, AWS Data Engineering, or similar. Contributions to open source data or AI/ML projects. Location & Travel Location: Remote, Brazil Travel: Minimal; occasional travel for team meetings or training Language Fluent English required (written and verbal) Why Rimini Street? We are looking for talented, passionate people to help us build our future at Rimini Street. We hire only the best, the most extraordinary professionals and provide compensation, bonuses, and benefits to match the skills of our top-performing team members. Do you thrive in a fast-paced environment, enjoy growing together, and get excited about learning new skills? Are you looking for an opportunity to make a true impact as part of a team of extraordinary professionals? This is the place for you. Our work is challenging and meaningful. We start and end each day with a sense of achievement and purpose guided by our core values, the Four Cs: Company We dream big and innovate boldly. Colleagues We work with extraordinary people who create a culture of mutual respect and collaboration. Clients We relentlessly pursue solutions that help clients achieve their goals. Our unmatched client care is rooted in our passion for exceptional service. Community We believe in leaving the world a better place than we found it. With the Rimini Street Foundation, we’ve made positive impacts in six continents for over 425 charities. Accelerating Company Growth Nasdaq-listed under ticker symbol RMNI since October 2017 Over 6,000 signed clients, including over 200 of the Fortune 500 and Global 100 companies Over 2,000 team members, 30 offices in 21 countries, supporting clients in over 160 countries US and international recognition for industry leadership and philanthropic efforts Rimini Street is committed to creating a diverse and inclusive environment and is proud to be an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to age, race, color, religion, national origin, sexual orientation, gender or gender identity, disability, protected veteran status, or any other characteristic protected by law. Unsolicited resumes will be ineligible for referral fees. In 2005, we disrupted the industry by pioneering extraordinary enterprise software support powered by extraordinary people. Our work is challenging and meaningful, allowing everyone at Rimini Street to start and end each day with a sense of achievement and purpose. We are looking for talented, passionate people to help us build our future at Rimini Street. Feel free to connect with us by signing up to receive updates on new jobs and opportunities! Create an account and sign up to create your job alerts! Link below: Rimini Street Career Job Alerts
Чем предстоит заниматься
- Design and build RAG pipelines that retrieve and rerank relevant context for AI agents.
- Implement and manage embedding and vector infrastructure, including pgvector-based storage and index optimization.
- Build data ingestion and ETL/processing pipelines from Salesforce, ServiceNow, documentation, email archives, and ERP logs, with data quality monitoring.
- Design knowledge architecture: taxonomies, knowledge graphs, versioning, and feedback loops for retrieval improvement.
- Integrate cloud data platform (Snowflake) with vector stores and design efficient synchronization and data access patterns.
Что требуется
- 5+ years of data engineering experience.
- 1–2+ years focused on AI/ML data pipelines or RAG systems.
- Hands-on experience with vector databases, embeddings, and similarity search (pgvector or Pinecone/Qdrant/Weaviate).
- Production experience with PostgreSQL and strong SQL; Python for data processing and pipelines.
- Experience building ETL/ELT pipelines and integrating data from diverse source systems (Salesforce, ServiceNow).
Rimini Street, Inc. (Nasdaq: RMNI), a Russell 2000® Company, is a proven, trusted global provider of end-to-end, mission-critical enterprise software support, managed services and innovative Agentic AI ERP solutions, and is the leading third-party support provider for Oracle, SAP and VMware software. The Company has signed thousands of IT service contracts with Fortune Global 100, Fortune 500, midmarket, public sector and government organizations who have leveraged the Rimini Smart Path™ methodology to achieve better operational outcomes, billions of US dollars in savings and fund AI and other innovation.
Что говорят о компании
3.0/ 5
- Flexible work hours and remote work options are appreciated by many employees.
- Some employees enjoy the collaborative and supportive team culture.
- Concerns about management effectiveness and communication have been raised.