Applied AI Data Engineer (Vector Databases, Data Management)

Remote Full-time
About the position As an Applied AI Data Engineer, you will be responsible for building data pipelines, vector embeddings, and retrieval mechanisms that power AI reasoning systems. Your work ensures that LLMs remain grounded in fact, efficiently retrieving high-quality, contextually relevant data without noise or hallucinations. You will design and implement features that harness vector search, retrieval-augmented generation (RAG), and domain-specific embeddings, directly influencing how AI models store, retrieve, and apply knowledge at scale. Responsibilities • Build and optimize data pipelines that transform incoming documents into high-quality embeddings for AI retrieval. • Design and implement vector search strategies using Pinecone, Weaviate, FAISS, or Vespa to improve AI response relevance. • Develop retrieval-augmented generation (RAG) workflows, ensuring models access up-to-date and high-quality context. • Fine-tune chunking strategies and indexing frequencies to enhance information recall and factual accuracy. • Integrate hybrid search approaches (semantic + keyword) to improve precision and efficiency in knowledge retrieval. • Monitor retrieval logs and LLM interaction patterns, adjusting embedding configurations for maximum relevance. • Compare model performance (GPT-4, Claude, Llama 2) across different embedding structures and refine tuning strategies. • Experiment with metadata filtering techniques to dynamically surface the most relevant data for AI reasoning agents. • Collaborate with ML engineers and AI researchers to ensure data pipelines align with evolving AI capabilities. Requirements • 5-8+ years of experience in Data Engineering, AI Systems, or Machine Learning Infrastructure. • 3+ years of hands-on experience working with vector databases, embeddings, and retrieval-augmented generation (RAG). • Strong understanding of vector search algorithms, indexing strategies, and hybrid search techniques. • Expertise in building and scaling data pipelines for AI-driven applications. • Proficiency in Python, along with experience using libraries such as Hugging Face, LangChain, and OpenAI SDKs. • Hands-on experience with vector database platforms (Pinecone, Weaviate, FAISS, ChromaDB, or Vespa). • Deep knowledge of LLM retrieval strategies, chunking methodologies, and context optimization. • Familiarity with semantic search, keyword search, and metadata filtering techniques. • Strong grasp of data governance, security, and optimization for AI-driven knowledge retrieval. • Experience integrating retrieval mechanisms with multi-agent AI systems. Nice-to-haves • Experience in fine-tuning transformer models for domain-specific retrieval tasks. • Familiarity with real-time indexing and adaptive embedding refresh strategies. • Understanding of LLM hallucination mitigation and factual consistency techniques. • Experience building scalable knowledge graphs and structured AI databases. • Background in AI-powered document processing and knowledge extraction. Benefits • Medical, dental, vision insurance. • Short and long-term disability insurance. • Life insurance. • 401k available on the first day of the month after start date. • Flexible PTO. Apply tot his job
Apply Now

Similar Opportunities

Experienced Registered Behavior Technician for In-Home ABA Therapy - Atlanta, GA

Remote Full-time

Immediate Hiring: Experienced Registered Behavioral Technician (RBT) for Clinic-Based ABA Therapy Services

Remote Full-time

Experienced Registered Behavioral Technician (RBT) - ABA Therapy for Children with Autism Spectrum Disorder

Remote Full-time

Experienced Registered Nurse - Telehealth: Providing Remote Care Coordination and Patient Support

Remote Full-time

Experienced Substitute Teacher for Riverside County Schools - Join Scoot Education's Innovative Team

Remote Full-time

Experienced Substitute Teacher for San Bernardino County - Flexible Schedules & Competitive Pay

Remote Full-time

Experienced School Year Instructional Coach for High-Dosage Tutoring Programs in Edgewater Park, NJ

Remote Full-time

Experienced School Year Tutor for K-8 Students in Math and Literacy - Mickleton, NJ

Remote Full-time

Experienced Secondary Social Studies Teacher for Kansas - Flexible Hybrid Remote Arrangement

Remote Full-time

USPS Office Helper

Remote Full-time

[Remote/WFM] Amazon Remote Work From Home Jobs - Hiring Now

Remote Full-time

Freelance Jr. Designer

Remote Full-time

Junior Crypto Analyst & Trader (Remote, Training Included)

Remote Full-time

Experienced Customer Support Representative – Home Advisor for Innovative Technology Solutions at blithequark

Remote Full-time

Experienced Data Entry Remote Specialist – Join the Innovative Team at Apple and Shape the Future of Technology from the Comfort of Your Own Home

Remote Full-time

senior product manager, Beverage US Licensed Stores (Remote-US)

Remote Full-time

Vertex Summer Intern 2026, Clinical and Quantitative Pharmacology

Remote Full-time

Experienced Gymnastics Instructor for Children's Development Programs in San Antonio, TX

Remote Full-time

**Experienced Part-Time Remote Data Entry Specialist – Entry-Level Opportunity with blithequark**

Remote Full-time

Client Success Representative - Overnight Shift - Remote Opportunity with NewtekOne

Remote Full-time
← Back to Home