Data Engineer
Quikr
full-remotemidpermanentbackenddata United States 45 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
PythonSQLApache SparkAWSData PipelinesDistributed Data ProcessingData QualityOrchestrationNoSQLMachine Learning Workflows
About the role
Role Overview
Data Engineer to build and scale data infrastructure that supports AI-driven products and research. You will design distributed data pipelines, manage large-scale datasets in the cloud, and enable reliable data systems for processing, experimentation, analytics, and machine learning workflows.
Responsibilities
- Design, build, and maintain scalable data pipelines for ingesting, processing, transforming, and distributing large datasets from multiple sources
- Develop and optimize distributed data processing workflows using Apache Spark and cloud-native technologies
- Design and maintain scalable data storage solutions across SQL and NoSQL databases
- Build data architectures on AWS for high-volume ingestion, processing, storage, and distribution
- Write efficient, maintainable Python and SQL for extraction, transformation, validation, and analysis
- Implement data quality, validation, monitoring, and reliability practices across pipelines and storage systems
- Optimize workflows for performance, scalability, reliability, and cost efficiency
- Implement automation and orchestration to support repeatable data operations
- Collaborate with AI researchers, data scientists, and software engineers to support data-intensive applications and experimentation
- Troubleshoot pipeline and data infrastructure issues and implement corrective measures
Required Qualifications
- Strong proficiency in Python and SQL
- Hands-on experience with Apache Spark (or other distributed processing frameworks)
- Experience designing and operating data pipelines in AWS (or comparable cloud environments)
- Experience working with both SQL and NoSQL databases
- Experience processing and managing large-scale datasets in distributed environments
- Strong understanding of data partitioning, distributed processing, performance optimization, and scalable data architecture
- Experience implementing data quality, monitoring, and reliability practices
- Strong analytical and problem-solving skills; effective collaboration in a technical environment
Preferred Qualifications
- Experience supporting AI/ML workflows, data science initiatives, or research environments
- Familiarity with datasets and infrastructure used for model training, evaluation, or experimentation
- Experience with LLM-related data workflows (training/evaluation datasets, prompt experimentation)
- Experience with data visualization tools (e.g., Matplotlib, Seaborn, Plotly)
- Familiarity with orchestration/workflow automation and cloud-native data platforms
- Experience with high-volume or real-time data processing systems
About Quikr
Quikr is an online classifieds and digital marketplace company operating in the technology and consumer services space. The company supports data-driven products and research initiatives through technology teams that build and scale data infrastructure.
Scraped 8/9/2026