Sat.Sep 09, 2023 - Fri.Sep 15, 2023

article thumbnail

An Overview Of The Sate Of Data Orchestration In An Increasingly Complex Data Ecosystem

Data Engineering Podcast

Summary Data systems are inherently complex and often require integration of multiple technologies. Orchestrators are centralized utilities that control the execution and sequencing of interdependent operations. This offers a single location for managing visibility and error handling so that data platform engineers can manage complexity. In this episode Nick Schrock, creator of Dagster, shares his perspective on the state of data orchestration technology and its application to help inform its im

BI 208
article thumbnail

The Role of DevOps and CI/CD in Data Engineering

Confessions of a Data Guy

In the vast world of data, it’s not just about gathering and analyzing information anymore; it’s also about ensuring that data pipelines, processes, and platforms run seamlessly and efficiently. Nothing screams “why are flying by night,” than coming into a Data Team only to find no tests, no docs, no deployments, no Docker, no nothing. […] The post The Role of DevOps and CI/CD in Data Engineering appeared first on Confessions of a Data Guy.

Insiders

Sign Up for our Newsletter

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

article thumbnail

Data Management Principles for Data Science

KDnuggets

Back to Basics: Understanding key data management principles that data scientists should know.

article thumbnail

Best Practices for LLM Evaluation of RAG Applications

databricks

Chatbots are the most widely adopted use case for leveraging the powerful chat and reasoning capabilities of large language models (LLM). The retrieval.

article thumbnail

A Guide to Debugging Apache Airflow® DAGs

In Airflow, DAGs (your data pipelines) support nearly every use case. As these workflows grow in complexity and scale, efficiently identifying and resolving issues becomes a critical skill for every data engineer. This is a comprehensive guide with best practices and examples to debugging Airflow DAGs. You’ll learn how to: Create a standardized process for debugging to quickly diagnose errors in your DAGs Identify common issues with DAGs, tasks, and connections Distinguish between Airflow-relate

article thumbnail

Meta Quest 2: Defense through offense

Engineering at Meta

Meta’s Native Assurance team regularly performs manual code reviews as part of our ongoing commitment to improve the security posture of Meta’s products. In 2021, we discovered a vulnerability in the Meta Quest 2’s Android-based OS that never made it to production but helped us find new ways to improve the security of Meta Quest products. We’re sharing our journey to get arbitrary native code execution in the privileged VR Runtime service on the Meta Quest 2 by exploiting a memory corruption v

Bytes 138
article thumbnail

Data News — Week 23.37

Christophe Blefari

Facing the News ( credits ) Hello Data News readers. I'm still struggling to get back into my usual work rhythm. If you add the fact that last week I came up with fewer articles than I expected, this has led me to another blank page. Anyway, after 2 years of work, I have to accept and let go when necessary. But don't worry I don't forget you.

More Trending

article thumbnail

Introducing MLflow 2.7 with new LLMOps capabilities

databricks

As part of MLflow 2’s support for LLMOps, we are excited to introduce the latest updates to support prompt engineering in MLflow 2.7. A.

article thumbnail

Apache Flink best practices - Flink Forward lessons learned

Waitingforcode

I won't hide it, I'm still a fresher in the Apache Flink world and despite my past streaming experiences with Apache Spark Structured Streaming and GCP Dataflow, I need to learn. And to learn a new tool or concept, there is nothing better than watching some conference talks!

IT 130
article thumbnail

How To Set Up Your Data Analytics Team For Success – Centralized vs Decentralized vs Federated Data Teams

Seattle Data Guy

Success in the data world hinges on team setup. I’ve delved into onboarding and standards in previous articles, but never into the structure of data teams. Typically, there are three configurations: Centralized, Decentralized, and Federated. Most companies I’ve seen use a mix of these. While the newest tech breakthroughs grab headlines, team organization is the… Read more The post How To Set Up Your Data Analytics Team For Success – Centralized vs Decentralized vs Federat

article thumbnail

Statistics in Data Science: Theory and Overview

KDnuggets

High-level exploration of the role of statistics in data science.

article thumbnail

Mastering Apache Airflow® 3.0: What’s New (and What’s Next) for Data Orchestration

Speaker: Tamara Fingerlin, Developer Advocate

Apache Airflow® 3.0, the most anticipated Airflow release yet, officially launched this April. As the de facto standard for data orchestration, Airflow is trusted by over 77,000 organizations to power everything from advanced analytics to production AI and MLOps. With the 3.0 release, the top-requested features from the community were delivered, including a revamped UI for easier navigation, stronger security, and greater flexibility to run tasks anywhere at any time.

article thumbnail

Power Holistic Customer Insights with Salesforce and Snowflake Data Sharing-Based Integration

Snowflake

Snowflake and Salesforce have built on our existing partnership to unify the full breadth of customer and business data and generate actionable insights for our customers. We are happy to announce the general availability of Bring Your Own Lake (BYOL) Data Sharing with the Snowflake Data Cloud from Salesforce Data Cloud. Organizations can now leverage Salesforce data directly in Snowflake via zero-ETL data sharing to accelerate decision-making and help streamline business processes.

Finance 129
article thumbnail

A Watershed Moment

ArcGIS

Updated data from the Watershed Boundary Dataset (WBD) are added to Living Atlas as new feature services.

Datasets 128
article thumbnail

The Power of a Trusted Data Lakehouse: Go Bust or Boom

databricks

Special thanks to our partners at Immuta, Alation, and Anomalo for their collaboration on the content and technical assets from this article. The.

Data 118
article thumbnail

Linear Regression from Scratch with NumPy

KDnuggets

Mastering the Basics of Linear Regression and Fundamentals of Gradient Descent and Loss Minimization.

article thumbnail

Agent Tooling: Connecting AI to Your Tools, Systems & Data

Speaker: Alex Salazar, CEO & Co-Founder @ Arcade | Nate Barbettini, Founding Engineer @ Arcade | Tony Karrer, Founder & CTO @ Aggregage

There’s a lot of noise surrounding the ability of AI agents to connect to your tools, systems and data. But building an AI application into a reliable, secure workflow agent isn’t as simple as plugging in an API. As an engineering leader, it can be challenging to make sense of this evolving landscape, but agent tooling provides such high value that it’s critical we figure out how to move forward.

article thumbnail

How Marriott Modernized Their Data Architecture with Snowflake

Snowflake

More than 50% of data leaders recently surveyed by BCG said the complexity of their data architecture is a significant pain point in their enterprise. Companies hampered by legacy data architectures are often plagued by a high total cost of ownership (TCO), an inability to govern data, and a lack of scalability as their data volumes grow. “As a result,” says BCG, “many companies find themselves at a tipping point, at risk of drowning in a deluge of data, overburdened with complexity and costs.

article thumbnail

Mode + ThoughtSpot recognized as Leaders in Snowflake’s 2023 Modern Marketing Data Stack awards

ThoughtSpot

We’re thrilled to announce that both ThoughtSpot and Mode ( acquired by ThoughtSpot in July 2023 ) have been recognized as Leaders in Snowflake's recent Modern Marketing Data Stack report! Given the ever-evolving landscape of modern data analytics products, organizations are looking to ThoughtSpot and Mode when seeking innovative solutions—helping them harness the power of their marketing data.

article thumbnail

Improve Lakehouse Security Monitoring using System Tables in Databricks Unity Catalog

databricks

As the lakehouse becomes increasingly mission-critical to data-forward organizations, so too grows the risk that unexpected events, outages, and security incidents may derail.

Systems 110
article thumbnail

10 Math Concepts for Programmers

KDnuggets

The not so secret behind becoming a proficient programmer - Math & it’s top 10 concepts.

article thumbnail

How to Modernize Manufacturing Without Losing Control

Speaker: Andrew Skoog, Founder of MachinistX & President of Hexis Representatives

Manufacturing is evolving, and the right technology can empower—not replace—your workforce. Smart automation and AI-driven software are revolutionizing decision-making, optimizing processes, and improving efficiency. But how do you implement these tools with confidence and ensure they complement human expertise rather than override it? Join industry expert Andrew Skoog as he explores how manufacturers can leverage automation to enhance operations, streamline workflows, and make smarter, data-dri

article thumbnail

6 Tips for Setting the Price of Your Data Product

Snowflake

Building your data product is only the beginning. You’ve considered a wide variety of use cases, and settled on the one you’ll focus on. Maybe you’re going to help hospitals predict emergency room visits and optimize their staffing. Or you’re going to enable restaurants to reduce their food waste. Or maybe you just have some really unique data that you think might be of use to someone.

Food 119
article thumbnail

Create trusted insights with Verified Liveboards

ThoughtSpot

ThoughtSpot users can easily create content with data using our intuitive, AI-powered search experience. However, business users sometimes find themselves asking a critical question: which content should I trust and use for my specific business use case? For example, if there are ten “Sales Performance” Liveboards created by different authors, you may wonder which is the golden version—the Liveboard that is reviewed, approved, and consistently maintained.

article thumbnail

Introducing Apache Spark™ 3.5

databricks

Today, we are happy to announce the availability of Apache Spark™ 3.5 on Databricks as part of Databricks Runtime 14.0. We extend our s.

article thumbnail

Understanding Machine Learning Algorithms: An In-Depth Overview

KDnuggets

Understanding Machine Learning: Exposing the Tasks, Algorithms, and Selecting the Best Model.

article thumbnail

The Ultimate Guide to Apache Airflow DAGS

With Airflow being the open-source standard for workflow orchestration, knowing how to write Airflow DAGs has become an essential skill for every data engineer. This eBook provides a comprehensive overview of DAG writing features with plenty of example code. You’ll learn how to: Understand the building blocks DAGs, combine them in complex pipelines, and schedule your DAG to run exactly when you want it to Write DAGs that adapt to your data at runtime and set up alerts and notifications Scale you

article thumbnail

Data Cloud Industry Day 2023: Your Event Guide

Snowflake

The first annual Data Cloud Industry Day is here! Data Cloud Industry Day 2023 is a free virtual event on September 28, 2023, dedicated to what’s possible for you and your industry in the world of data. From leading-edge innovations to seamless solutions to your toughest industry-specific challenges, Industry Day provides the insights and information you need to drive business value with your data.

Cloud 111
article thumbnail

Last Mile Data Processing with Ray

Pinterest Engineering

Raymond Lee | Software Engineer II; Qingxian Lai | Sr. Software Engineer; Karthik Anantha Padmanabhan | Manager II, Engineering; Se Won Jang | Manager II, Engineering Photo by Claudio Schwarz on Unsplash Our mission at Pinterest is to bring everyone the inspiration to create the life they love. Machine Learning plays a crucial role in this mission. It allows us to continuously deliver high-quality inspiration to our 460 million monthly active users, curated from billions of pins on our platform.

article thumbnail

Crossing Bridges: Reporting on NYC taxi data with RStudio and Databricks

databricks

As data enthusiasts, we love uncovering stories in datasets. With Posit’s RStudio Desktop and Databricks Lakehouse, you can analyze data with dplyr, create i.

Datasets 105
article thumbnail

Getting Started with SQL in 5 Steps

KDnuggets

This comprehensive SQL tutorial covers everything from setting up your SQL environment to mastering advanced concepts like joins, subqueries, and optimizing query performance. With step-by-step examples, this guide is perfect for beginners looking to enhance their data management skills.

SQL 145
article thumbnail

Apache Airflow® Best Practices: DAG Writing

Speaker: Tamara Fingerlin, Developer Advocate

In this new webinar, Tamara Fingerlin, Developer Advocate, will walk you through many Airflow best practices and advanced features that can help you make your pipelines more manageable, adaptive, and robust. She'll focus on how to write best-in-class Airflow DAGs using the latest Airflow features like dynamic task mapping and data-driven scheduling!

article thumbnail

Marketing Success in the Age of AI Requires a Modern Marketing Data Stack

Snowflake

Data is essential to marketing. It’s how we know our audience and measure campaign outcomes. It shows us where to adjust a campaign on the fly, for even better results. But working with data is increasingly complex, and having the right stack of technologies is invaluable. To help marketers understand the rapidly changing world of data technology, we’ve just released our second annual Modern Marketing Data Stack report.

Media 105
article thumbnail

Path Representation in Python

Towards Data Science

Here’s why you should avoid representing paths as strings and use Pathlib instead Continue reading on Towards Data Science »

Python 98
article thumbnail

New Looker + ThoughtSpot Connector: Where semantic modeling meets natural language search

ThoughtSpot

Semantic layers are a game changer, allowing organizations to define metrics and business logic in one, centralized location. Because business users can trust that their data is built on a single source of truth, the semantic layer also empowers self-service analytics. Looker Modeler has become a leader among semantic layers, allowing users to seamlessly layer on top of their business data.

BI 98
article thumbnail

KDnuggets News, September 13: Getting Started with SQL in 5 Steps • Introduction to Databases in Data Science

KDnuggets

Getting Started with SQL in 5 Steps • Introduction to Databases in Data Science • Time 100 AI: The Most Influential?

article thumbnail

How to Achieve High-Accuracy Results When Using LLMs

Speaker: Ben Epstein, Stealth Founder & CTO | Tony Karrer, Founder & CTO, Aggregage

When tasked with building a fundamentally new product line with deeper insights than previously achievable for a high-value client, Ben Epstein and his team faced a significant challenge: how to harness LLMs to produce consistent, high-accuracy outputs at scale. In this new session, Ben will share how he and his team engineered a system (based on proven software engineering approaches) that employs reproducible test variations (via temperature 0 and fixed seeds), and enables non-LLM evaluation m