Sat.Sep 09, 2023 - Fri.Sep 15, 2023

article thumbnail

An Overview Of The Sate Of Data Orchestration In An Increasingly Complex Data Ecosystem

Data Engineering Podcast

Summary Data systems are inherently complex and often require integration of multiple technologies. Orchestrators are centralized utilities that control the execution and sequencing of interdependent operations. This offers a single location for managing visibility and error handling so that data platform engineers can manage complexity. In this episode Nick Schrock, creator of Dagster, shares his perspective on the state of data orchestration technology and its application to help inform its im

BI 208
article thumbnail

The Role of DevOps and CI/CD in Data Engineering

Confessions of a Data Guy

In the vast world of data, it’s not just about gathering and analyzing information anymore; it’s also about ensuring that data pipelines, processes, and platforms run seamlessly and efficiently. Nothing screams “why are flying by night,” than coming into a Data Team only to find no tests, no docs, no deployments, no Docker, no nothing. […] The post The Role of DevOps and CI/CD in Data Engineering appeared first on Confessions of a Data Guy.

Insiders

Sign Up for our Newsletter

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

article thumbnail

GPT and LLMs from a Data Engineering Perspective

Jesse Anderson

There has been quite a bit of writing covering GPT and LLMs from data science and business perspectives. I haven’t seen much from the data engineering side. Let me share my perspective, having been in data and AI for a while and using LLMs before they became popular. It is interesting to see the general public having the same amount of excitement as there was a year ago in the LLM space.

article thumbnail

Data News — Week 23.37

Christophe Blefari

Facing the News ( credits ) Hello Data News readers. I'm still struggling to get back into my usual work rhythm. If you add the fact that last week I came up with fewer articles than I expected, this has led me to another blank page. Anyway, after 2 years of work, I have to accept and let go when necessary. But don't worry I don't forget you.

article thumbnail

Apache Airflow® Best Practices for ETL and ELT Pipelines

Whether you’re creating complex dashboards or fine-tuning large language models, your data must be extracted, transformed, and loaded. ETL and ELT pipelines form the foundation of any data product, and Airflow is the open-source data orchestrator specifically designed for moving and transforming data in ETL and ELT pipelines. This eBook covers: An overview of ETL vs.

article thumbnail

Apache Flink best practices - Flink Forward lessons learned

Waitingforcode

I won't hide it, I'm still a fresher in the Apache Flink world and despite my past streaming experiences with Apache Spark Structured Streaming and GCP Dataflow, I need to learn. And to learn a new tool or concept, there is nothing better than watching some conference talks!

IT 130
article thumbnail

How To Set Up Your Data Analytics Team For Success – Centralized vs Decentralized vs Federated Data Teams

Seattle Data Guy

Success in the data world hinges on team setup. I’ve delved into onboarding and standards in previous articles, but never into the structure of data teams. Typically, there are three configurations: Centralized, Decentralized, and Federated. Most companies I’ve seen use a mix of these. While the newest tech breakthroughs grab headlines, team organization is the… Read more The post How To Set Up Your Data Analytics Team For Success – Centralized vs Decentralized vs Federat

More Trending

article thumbnail

Best Practices for LLM Evaluation of RAG Applications

databricks

Chatbots are the most widely adopted use case for leveraging the powerful chat and reasoning capabilities of large language models (LLM). The retrieval.

article thumbnail

Meta Quest 2: Defense through offense

Engineering at Meta

Meta’s Native Assurance team regularly performs manual code reviews as part of our ongoing commitment to improve the security posture of Meta’s products. In 2021, we discovered a vulnerability in the Meta Quest 2’s Android-based OS that never made it to production but helped us find new ways to improve the security of Meta Quest products. We’re sharing our journey to get arbitrary native code execution in the privileged VR Runtime service on the Meta Quest 2 by exploiting a memory corruption v

Bytes 132
article thumbnail

Mode + ThoughtSpot recognized as Leaders in Snowflake’s 2023 Modern Marketing Data Stack awards

ThoughtSpot

We’re thrilled to announce that both ThoughtSpot and Mode ( acquired by ThoughtSpot in July 2023 ) have been recognized as Leaders in Snowflake's recent Modern Marketing Data Stack report! Given the ever-evolving landscape of modern data analytics products, organizations are looking to ThoughtSpot and Mode when seeking innovative solutions—helping them harness the power of their marketing data.

article thumbnail

The 5 Best AI Tools For Maximizing Productivity

KDnuggets

KDnuggets reviews a diverse set of 5 AI tools to help maximize your productivity. Have a look and see what our recommendations include.

136
136
article thumbnail

Apache Airflow®: The Ultimate Guide to DAG Writing

Speaker: Tamara Fingerlin, Developer Advocate

In this new webinar, Tamara Fingerlin, Developer Advocate, will walk you through many Airflow best practices and advanced features that can help you make your pipelines more manageable, adaptive, and robust. She'll focus on how to write best-in-class Airflow DAGs using the latest Airflow features like dynamic task mapping and data-driven scheduling!

article thumbnail

Power Holistic Customer Insights with Salesforce and Snowflake Data Sharing-Based Integration

Snowflake

Snowflake and Salesforce have built on our existing partnership to unify the full breadth of customer and business data and generate actionable insights for our customers. We are happy to announce the general availability of Bring Your Own Lake (BYOL) Data Sharing with the Snowflake Data Cloud from Salesforce Data Cloud. Organizations can now leverage Salesforce data directly in Snowflake via zero-ETL data sharing to accelerate decision-making and help streamline business processes.

Finance 108
article thumbnail

Last Mile Data Processing with Ray

Pinterest Engineering

Raymond Lee | Software Engineer II; Qingxian Lai | Sr. Software Engineer; Karthik Anantha Padmanabhan | Manager II, Engineering; Se Won Jang | Manager II, Engineering Photo by Claudio Schwarz on Unsplash Our mission at Pinterest is to bring everyone the inspiration to create the life they love. Machine Learning plays a crucial role in this mission. It allows us to continuously deliver high-quality inspiration to our 460 million monthly active users, curated from billions of pins on our platform.

article thumbnail

Introducing MLflow 2.7 with new LLMOps capabilities

databricks

As part of MLflow 2’s support for LLMOps, we are excited to introduce the latest updates to support prompt engineering in MLflow 2.7. A.

article thumbnail

KDnuggets News, September 13: Getting Started with SQL in 5 Steps • Introduction to Databases in Data Science

KDnuggets

Getting Started with SQL in 5 Steps • Introduction to Databases in Data Science • Time 100 AI: The Most Influential?

article thumbnail

Optimizing The Modern Developer Experience with Coder

Many software teams have migrated their testing and production workloads to the cloud, yet development environments often remain tied to outdated local setups, limiting efficiency and growth. This is where Coder comes in. In our 101 Coder webinar, you’ll explore how cloud-based development environments can unlock new levels of productivity. Discover how to transition from local setups to a secure, cloud-powered ecosystem with ease.

article thumbnail

6 Tips for Setting the Price of Your Data Product

Snowflake

Building your data product is only the beginning. You’ve considered a wide variety of use cases, and settled on the one you’ll focus on. Maybe you’re going to help hospitals predict emergency room visits and optimize their staffing. Or you’re going to enable restaurants to reduce their food waste. Or maybe you just have some really unique data that you think might be of use to someone.

Food 104
article thumbnail

A Watershed Moment

ArcGIS

Updated data from the Watershed Boundary Dataset (WBD) are added to Living Atlas as new feature services.

Datasets 127
article thumbnail

Measuring Technical Debt to Avoid the Boiling Frog Syndrome

Booking.com Engineering

source Software development is all about change. And, over the lifespan of our software, the goal is to implement required changes in a reasonable amount of time. Whether the changes are technical in nature, like an urgent security upgrade, or stem from a business need, such as building a new feature to make us more competitive in target markets — how fast we can change is critical.

Coding 93
article thumbnail

Getting Started with SQL in 5 Steps

KDnuggets

This comprehensive SQL tutorial covers everything from setting up your SQL environment to mastering advanced concepts like joins, subqueries, and optimizing query performance. With step-by-step examples, this guide is perfect for beginners looking to enhance their data management skills.

SQL 116
article thumbnail

15 Modern Use Cases for Enterprise Business Intelligence

Large enterprises face unique challenges in optimizing their Business Intelligence (BI) output due to the sheer scale and complexity of their operations. Unlike smaller organizations, where basic BI features and simple dashboards might suffice, enterprises must manage vast amounts of data from diverse sources. What are the top modern BI use cases for enterprise businesses to help you get a leg up on the competition?

article thumbnail

How Marriott Modernized Their Data Architecture with Snowflake

Snowflake

More than 50% of data leaders recently surveyed by BCG said the complexity of their data architecture is a significant pain point in their enterprise. Companies hampered by legacy data architectures are often plagued by a high total cost of ownership (TCO), an inability to govern data, and a lack of scalability as their data volumes grow. “As a result,” says BCG, “many companies find themselves at a tipping point, at risk of drowning in a deluge of data, overburdened with complexity and costs.

article thumbnail

Crossing Bridges: Reporting on NYC taxi data with RStudio and Databricks

databricks

As data enthusiasts, we love uncovering stories in datasets. With Posit’s RStudio Desktop and Databricks Lakehouse, you can analyze data with dplyr, create i.

article thumbnail

How to Build an Interactive Real-Time Chat Application with Websockets?

Workfall

Reading Time: 11 minutes What is Socket.io? Socket.io , a widely-used JavaScript library, offers a framework for facilitating real-time, two-way communication between web clients (like browsers) and servers. It uses WebSockets as the primary communication method but also offers fallback options such as long polling for environments where WebSockets may not be supported.

article thumbnail

Closed Source VS Open Source Image Annotation

KDnuggets

This blog strikes a comparison between open-source and closed-source image annotation tools and how it makes the life of AI model developers easy and convenient.

IT 116
article thumbnail

Prepare Now: 2025s Must-Know Trends For Product And Data Leaders

Speaker: Jay Allardyce, Deepak Vittal, and Terrence Sheflin

As we look ahead to 2025, business intelligence and data analytics are set to play pivotal roles in shaping success. Organizations are already starting to face a host of transformative trends as the year comes to a close, including the integration of AI in data analytics, an increased emphasis on real-time data insights, and the growing importance of user experience in BI solutions.

article thumbnail

Data Cloud Industry Day 2023: Your Event Guide

Snowflake

The first annual Data Cloud Industry Day is here! Data Cloud Industry Day 2023 is a free virtual event on September 28, 2023, dedicated to what’s possible for you and your industry in the world of data. From leading-edge innovations to seamless solutions to your toughest industry-specific challenges, Industry Day provides the insights and information you need to drive business value with your data.

Cloud 93
article thumbnail

Introducing Apache Spark™ 3.5

databricks

Today, we are happy to announce the availability of Apache Spark™ 3.5 on Databricks as part of Databricks Runtime 14.0. We extend our s.

article thumbnail

Why Your Data Pipelines Need Closed-Loop Feedback Control

Towards Data Science

Realities of company and cloud complexities require new levels of control and autonomy to meet business goals at scale Image by Cosmin Paduraru As data teams scale up on the cloud, data platform teams need to ensure the workloads they are responsible for are meeting business objectives. At scale with dozens of data engineers building hundreds of production jobs, controlling their performance at scale is untenable for a myriad of reasons from technical to human.

article thumbnail

KDnuggets Top Posts for August 2023: Forget ChatGPT, This New AI Assistant Will Change the Way You Work

KDnuggets

Forget ChatGPT, This New AI Assistant Is Leagues Ahead and Will Change the Way You Work Forever • 7 Projects Built with Generative AI • Best Python Tools for Building Generative AI Applications Cheat Sheet • Harnessing ChatGPT for Automated Data Cleaning and Preprocessing • Data Scientists Need to Specialize to Survive the Tech Winter • 7 Steps to Mastering Data Cleaning and Preprocessing Techniques • 5 Ways You Can Use ChatGPT's Code Interpreter For Data Science • The Best Courses for AI from U

article thumbnail

Apache Airflow®: The Ultimate Guide to DAG Writing

Speaker: Tamara Fingerlin, Developer Advocate

In this new webinar, Tamara Fingerlin, Developer Advocate, will walk you through many Airflow best practices and advanced features that can help you make your pipelines more manageable, adaptive, and robust. She'll focus on how to write best-in-class Airflow DAGs using the latest Airflow features like dynamic task mapping and data-driven scheduling!

article thumbnail

Marketing Success in the Age of AI Requires a Modern Marketing Data Stack

Snowflake

Data is essential to marketing. It’s how we know our audience and measure campaign outcomes. It shows us where to adjust a campaign on the fly, for even better results. But working with data is increasingly complex, and having the right stack of technologies is invaluable. To help marketers understand the rapidly changing world of data technology, we’ve just released our second annual Modern Marketing Data Stack report.

Media 87
article thumbnail

Improve Lakehouse Security Monitoring using System Tables in Databricks Unity Catalog

databricks

As the lakehouse becomes increasingly mission-critical to data-forward organizations, so too grows the risk that unexpected events, outages, and security incidents may derail.

Systems 88
article thumbnail

A Talented Team, Innovative Technology, and The Opportunity to Grow. There Is No Place Like Cloudera

Cloudera

I started my current career path with Hortonworks in 2016, back when we still had to tell people what Hadoop was. Once I got to work with all the amazing open-source Apache tools I was hooked. I found Apache NiFi especially interesting. Soon after, I became a huge fan of Apache Kafka. Coupled with amazing technology was an amazing team that only grew and improved with the merger with Cloudera.

article thumbnail

Working with Big Data: Tools and Techniques

KDnuggets

Where do you start in a field as vast as big data? Which tools and techniques to use? We explore this and talk about the most common tools in big data.

article thumbnail

The Cloud Development Environment Adoption Report

Cloud Development Environments (CDEs) are changing how software teams work by moving development to the cloud. Our Cloud Development Environment Adoption Report gathers insights from 223 developers and business leaders, uncovering key trends in CDE adoption. With 66% of large organizations already using CDEs, these platforms are quickly becoming essential to modern development practices.