Position: Data Analyst / Data Engineer Intern
Company: DeHaat
Location: Gurugram, Haryana, India
Job type: Full-time Internship
Job mode: On-site
Job requisition id: Not specified in the original JD
Years of experience: 0-3 years (College students can also try)
Company description
DeHaat stands as one of the most rapidly expanding technology ventures within the Indian agricultural sector, pioneering full-stack solutions tailored specifically for rural farming communities.
Driven by advanced artificial intelligence and digital innovations, the organization actively restructures conventional agricultural ecosystems, optimizing end-to-end supply chain logistics and boosting farm productivity.
The company currently maintains an operational footprint across 12 prominent agricultural states, powering a ground-level network of over 15,000 localized DeHaat distribution centers and partnering with 503 Farmer Producer Organizations.
Through these digital hubs, the platform directly impacts and supports more than 12.8 million farmers across India by offering personalized crop advisories in native languages covering 30 distinct crops.
Established by visionaries from premier academic institutes such as IIT Delhi, IIT Kharagpur, and IIM Ahmedabad, the venture has built a highly scalable enterprise architecture.
As a backed startup experiencing exponential year-on-year growth, the firm continuously receives high-profile accolades from global institutions like the Bill Gates Foundation, NITI Aayog, NASSCOM, Forbes, and The Economic Times.
Profile overview
The Product & Technology group at DeHaat is recruiting a motivated Data Analyst / Data Engineer Intern to directly contribute to building, monitoring, and streamlining core data processing mechanisms.
In this technical position, the candidate will work closely alongside engineering leads and product managers to develop robust data pipelines, maintain dynamic dashboards, and manage operational reporting workflows.
The role requires strong foundational capabilities in modern database administration, pipeline infrastructure, and programmatic data transformation techniques.
A core focus lies in engineering robust data pipelines capable of transferring information across distributed cloud databases, internal platforms, APIs, and business applications.
The intern will be responsible for crafting, executing, and refining high-performance SQL queries designed for data extraction, operational reconciliation, and standard metric reporting.
Beyond traditional analytics, the candidate will build production-grade ETL and ELT processes utilizing Python scripts to perform validation routines, automation sequences, and complex data cleanses.
The selected engineer will monitor continuous automation jobs, identify system alerts, resolve failing workflows, and resolve schema mismatches across operational environments.
In addition to data infrastructure tasks, a major workload component revolves around developing, customizing, and scaling interactive management dashboards utilizing the Frappe Framework and ERPNext systems.
The candidate will interface backend database clusters directly into business dashboards, creating structured data models, customized key performance indicators, and dynamic search parameters.
This role acts as a bridge between pure software development and functional business requirements, granting deep exposure to enterprise product development and cloud architectures.
Qualifications
Demonstrated academic background or core practical knowledge in Computer Science, Data Engineering, Software Engineering, or related quantitative disciplines.
High proficiency in writing complex SQL queries involving multi-table joins, subqueries, Common Table Expressions, aggregations, and performance-tuned data manipulations.
Solid practical coding skills in Python tailored specifically for automation workflows, structural data transformation, and scripting tasks.
Fundamental mastery of relational database structures, normalized data modeling practices, and database schema organization.
Practical understanding of ETL/ELT pipeline creation, scheduled script execution, and data flow synchronization.
Advanced familiarity with spreadsheet tools like Google Sheets or Microsoft Excel for fast data verification and pivot analysis.
Solid baseline knowledge regarding RESTful API architectures, structural JSON object parsing, and API debugging via client utilities such as Postman.
Excellent analytical reasoning capabilities combined with systemic problem-solving skills to isolate root causes in failing pipelines.
Preferential consideration for applicants having hands-on experience or familiarity with the Frappe Framework or ERPNext platform.
Advantageous exposure to relational cloud database environments such as Amazon Web Services RDS, Amazon Redshift, PostgreSQL, or MySQL engine variants.
Familiarity with Linux server navigation, task schedulers, cron job management, and Git version control tools.
Baseline familiarity with web technologies including JavaScript, HTML, and CSS for quick frontend adjustments.
Active experience incorporating generative AI programming assistants, such as ChatGPT, Claude, or Cursor, into coding and troubleshooting tasks.
Additional info
The designated internship requires a dedicated on-site commitment duration ranging from 6 to 9 months at the corporate head office located in Gurugram, India.
Selected candidates must be prepared to join immediately upon receiving an official offer from the talent acquisition team.
The position operates entirely within an on-site office model, requiring active presence at the technology development facility in Haryana.
The intern will gain hands-on technical experience working directly on live, large-scale production environments containing real-world agricultural data datasets.
Excellent candidate growth path offering direct engagement with modern technologies including AWS infrastructure, Frappe application builds, and automated workflow orchestrations.
Candidates are expected to take complete technical ownership of assigned micro-projects, managing them from initial requirement collection through deployment and rigorous QA testing.
The role requires continuous cross-functional collaboration alongside engineering teams, business intelligence analysts, field operations, and product designers.
The company provides a technical environment that encourages developers to leverage AI-assisted tools for rapid prototyping, automated documentation creation, and standard operating procedure development.
Please click here to apply.

