Skip to main content
Neurons Lab

Data Engineer (UA/RU Language speaking)

RemotePoland only
Published
Role
Data Engineering
Experience
Mid
Employment
Part-time20h/week
Salary not disclosed
Check eligibility

Open to PL only. Set where you work from to check your eligibility.

No BS summary

Part-time Data Engineer with 4+ years in data engineering, strong Python and SQL, and real unstructured/semi-structured pipeline experience. Needs API integrations, entity resolution, orchestration, cloud data stack, PII/GDPR handling, and clear written English. UA/RU language speaking role in Poland.

Core skills

PythonUnstructured-data pipelinesEntity resolution

Required skills

SQLEmbeddingRetrieval infrastructureVector storespgvectorOpenSearchPineconeGraph storesAPI integrationGoogle WorkspaceM365SlackCRMRecord linkageAirflow/Step FunctionsAWS/GCPPII detectionRedactionEncryption

Optional skills

RAG

Required languages

UkrainianRussianEnglish Clear written

What you'll do

  • Stand up capture by default: notetaker on every call with speaker attribution, plus ingestion from mail, Slack and messengers — designed as opt-out, not opt-in, and reversible if the client changes their mind.
  • Backfill the archive: years of historical email, Slack, board protocols, decks and portfolio updates — parsed, deduplicated and dated correctly.
  • Build document parsing for the awkward long tail: PDFs, scanned board packs, spreadsheets, slide decks, forwarded attachments.
  • Implement identity / entity resolution: the same person across Slack handle, mail alias and calendar invite; the same portfolio company across a deck, a mail thread and a CRM record.
  • Build chunking and embedding pipelines and load the vector + graph stores behind the ontology the architect defines.
  • Implement incremental sync through the connector layer (MCP / Composio-class) — no full re-crawls, no silent drift, clear handling of edits and deletions.
  • Attach access scope and provenance to every record at ingestion, so permission-aware retrieval and audit are possible downstream rather than bolted on.
  • Run PII detection, redaction and retention logic; evidence to the client's security function what is stored, where, and for how long.
  • Orchestrate with Airflow / Step Functions; build repeatable, monitored pipelines rather than scripts, with alerting when a source stops flowing.
  • Keep cost and latency under control at volume — batching, incremental embedding, storage tiering — and report the unit economics.
  • Write runbooks so the client's own team can operate this after handover.

What they require

  • Strong Python and solid SQL
  • Unstructured-data pipelines: transcripts, mail, chat, documents — parsing, normalisation, deduplication
  • Embedding / retrieval infrastructure: chunking strategies, vector stores (pgvector, OpenSearch, Pinecone-class), plus loading a graph store
  • API and connector integration at scale: Google Workspace / M365, Slack, CRM; rate limits, pagination, incremental cursors, webhooks
  • Entity resolution / record linkage (deterministic + fuzzy) without a clean shared key
  • Orchestration: Airflow, Step Functions or equivalent; idempotent, restartable jobs
  • AWS and/or GCP data stack; comfortable in a private / VPC deployment
  • PII detection, redaction, encryption and retention in practice
  • Clear written English; documents for handover and works well async in a small distributed pod
  • GDPR applied to employee-generated data (mail, chat, meeting recordings) and EU data residency across multiple jurisdictions
  • Data lineage, provenance and audit patterns — and why an AI system needs them more, not less
  • How retrieval quality depends on ingestion quality — enough understanding of RAG to make the right upstream choices
  • Preferred: Well-Architected security and cost practice; awareness of financial-services expectations
  • 4+ years in data engineering, with real unstructured / semi-structured work (not only warehouse modelling)
  • Demonstrated experience integrating many third-party APIs into one coherent store, including historical backfill
  • Preferred: Experience building pipelines feeding an LLM / retrieval system
  • Experience handling sensitive personal data in a regulated or security-sensitive environment
  • Comfortable being the only data engineer on a small (2.5-FTE) pod, at part-time allocation, without hand-holding

Neurons Lab runs a group-wide AI Adoption Program for a major iGaming client: a holding of six game studios plus central business functions, 10+ companies, ~800–1,000 employees. The program combines business-team enablement, engineering enablement, and custom AI for game production.

AI
Salary not disclosed