Outmarket AI Logo

Outmarket AI

Data Engineer

Posted 9 Days Ago
Remote
Hiring Remotely in India
Junior
Remote
Hiring Remotely in India
Junior
Build and operate ingestion and transformation pipelines to turn messy insurance data into clean datasets. Design bronze/silver/fact layers, model and tune Postgres and analytical stores, define KPIs and reconciliation checks, orchestrate observable data jobs, and ensure end-to-end data quality for analytics and AI products.
The summary above was generated by AI

ABOUT OUTMARKET

Outmarket is the AI platform for insurance, trusted by more than 250 brokerages to run the work their business depends on. Commercial insurance still runs on dense documents and slow, manual workflows, and that is exactly what we automate: quote comparisons, coverage gap and tower analysis, policy review, and proposal generation, all grounded in our customers’ own data and source-cited so teams can trust the output.

The impact is concrete. Teams save 12 to 15 hours per person every week, cut errors by roughly 65 percent, and win more business, all on infrastructure that is SOC 2 Type II certified, single-tenant, and never used to train AI models. We are an AI-first company in both what we build and how we work, shipping quickly and in close partnership with the agencies that rely on us.

WHAT YOU’LL GET

  • A high-impact role with ownership from day one.

  • Competitive compensation and meaningful equity.

  • Direct collaboration with founders and real users.

  • Remote-first flexibility.

  • The opportunity to help build an AI-native product from the ground up.

ABOUT THE ROLE
We are hiring a Data Engineer to own the data plane that powers our analytics and AI products, from ingesting messy, real-world insurance data to the transformation pipelines and serving layer our customers rely on. You will turn fragmented source-system data into clean, trustworthy datasets that drive insights and automation.

WHY THIS ROLE

  • Own data infrastructure that is core to the product, not a side system.

  • Work with genuinely hard, messy, high-value insurance data.

  • Build the pipelines and models that everything analytical and AI-driven depends on.

WHAT YOU’LL DO

  • Build and operate ingestion pipelines from agency management systems, carriers, and documents.

  • Design transformation layers (bronze → silver → fact) and the models that serve insights and dashboards.

  • Model and tune data across Postgres and ClickHouse for correctness, performance, and cost.

  • Define dataset configurations, KPIs, and reconciliation checks that keep numbers trustworthy.

  • Orchestrate reliable, observable data jobs and own data quality end to end.

WHAT WE’RE LOOKING FOR

  • 2+ years building production data pipelines with strong SQL and Python.

  • Solid relational modeling experience (Postgres) and comfort with large analytical datasets.

  • A track record of turning messy, inconsistent source data into reliable product capabilities.

  • Strong ownership and comfort operating in a fast-moving environment.

BONUS IF YOU HAVE

  • Experience with ClickHouse or another columnar/analytical store.

  • Experience with workflow orchestration (e.g., Temporal, Airflow, Dagster).

  • Exposure to insurance, fintech, or other document- and data-heavy domains.

Similar Jobs

Yesterday
Remote or Hybrid
Senior level
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Design and operate scalable batch and streaming data platforms supporting machine learning and generative AI. Build pipelines for structured, unstructured, OCR, document, and image data; develop RAG, semantic search, and LLM-powered solutions; and establish data quality, observability, governance, orchestration, and deployment practices. Partner with stakeholders, lead platform scalability and cost optimization, mentor engineers, and translate ambiguous needs into production-ready technical roadmaps while securely handling sensitive data.
Top Skills: AirflowAmazon KinesisSparkAWSAzureAzure Event HubsChart.JsDatabricksDeequDelta LakeDockerGithub ActionsGCPGreat ExpectationsJavaKafkaKubernetesLlmsMlopsPlotlyPysparkPythonRagScalaSeabornSnowflakeSQLTerraform
2 Days Ago
Remote
Mid level
Mid level
Artificial Intelligence • Information Technology • Professional Services • Software • Analytics • Generative AI • Big Data Analytics
Configure and maintain Adobe Experience Platform and Real-Time CDP data pipelines, XDM schemas, datasets, identity resolution, Profile enablement, segmentation readiness, and destination activation. Validate data quality, troubleshoot ingestion and identity issues, manage sandboxes, and implement privacy, consent, governance, and access controls. Partner with architects, data engineers, consultants, analysts, data scientists, and marketing teams to support reporting, personalization, and machine learning readiness.
Top Skills: Adobe AnalyticsAdobe Experience PlatformAdobe I/O RuntimeAdobe Journey OptimizerAdobe Real-Time CdpAdobe Source ConnectorsAdobe TargetAPIsAWSAzureBatch IngestionBigQueryCcpaData PrepEltETLGCPGdprJavaScriptPythonQuery ServiceRedshiftSalesforce CdpSegmentSnowflakeSQLStreaming IngestionWeb SdkXdm
8 Days Ago
In-Office or Remote
Senior level
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead cloud data modernization by designing and implementing Azure/Snowflake/Databricks data platforms, building scalable ETL/ELT pipelines, ensuring data quality, security and governance, implementing CI/CD, mentoring engineers, and supporting healthcare data solutions and compliance.
Top Skills: AirflowAzureAzure Data Factory (Adf)Azure Data Lake Storage Gen2 (Adls Gen2)Change Data Capture (Cdc)CptDatabricksDatabricks GenieEtl/EltFacetsFhirGitGithub ActionsGithub CopilotHcpcsHl7Icd-10LlmsLoincPrompt EngineeringPysparkPythonRag PipelinesSnowflakeSnowflake CortexSparkSQL ServerSsisVector StoresVisioX12 Edi (837/835/834)

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account