Sat.Dec 24, 2022 - Fri.Dec 30, 2022

article thumbnail

Data News — must-read 2022 articles

Christophe Blefari

kitsch moment, from me to you ( credits ) Hey you, this is the last article of the year and it's gonna be about the articles and trends that made 2022 according to me. You'll see articles that I've already share during the year. 💡 You can also read the 2021's must-read that I've done one year and half ago or how to learn data engineering that contains key articles to understand the field.

article thumbnail

I asked ChatGPT to write a blog post about Data Engineering. Here it is.

Confessions of a Data Guy

Data engineering is a vital field within the realm of data science that focuses on the practical aspects of collecting, storing, and processing large amounts of data. It involves designing and building the infrastructure to store and process data, as well as developing the tools and systems to extract valuable insights and knowledge from that […] The post I asked ChatGPT to write a blog post about Data Engineering.

Insiders

Sign Up for our Newsletter

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

article thumbnail

Should We Get Rid Of ETLs?

Seattle Data Guy

AWS has jumped on the bandwagon of removing the need for ETLs. Snowflake announced this both with their hybrid tables and their partnership with Salesforce. Now, I do take a little issue with the naming “Zero ETLs”. Because at the very surface the functionality described is often closer to a zero integration future, which probably… Read more The post Should We Get Rid Of ETLs?

AWS 130
article thumbnail

Using Product Driven Development To Improve The Productivity And Effectiveness Of Your Data Teams

Data Engineering Podcast

Summary With all of the messaging about treating data as a product it is becoming difficult to know what that even means. Vishal Singh is the head of products at Starburst which means that he has to spend all of his time thinking and talking about the details of product thinking and its application to data. In this episode he shares his thoughts on the strategic and tactical elements of moving your work as a data professional from being task-oriented to being product-oriented and the long term i

Data Lake 130
article thumbnail

15 Modern Use Cases for Enterprise Business Intelligence

Large enterprises face unique challenges in optimizing their Business Intelligence (BI) output due to the sheer scale and complexity of their operations. Unlike smaller organizations, where basic BI features and simple dashboards might suffice, enterprises must manage vast amounts of data from diverse sources. What are the top modern BI use cases for enterprise businesses to help you get a leg up on the competition?

article thumbnail

Data Science Minimum: 10 Essential Skills You Need to Know to Start Doing Data Science

KDnuggets

Data science is ever-evolving, so mastering its foundational technical and soft skills will help you be successful in a career as a Data Scientist, as well as pursue advance concepts, such as deep learning and artificial intelligence.

article thumbnail

What is Apache Arrow? Asking for a friend.

Confessions of a Data Guy

We’ve all been in that spot, especially in tech. You wanted to fit in, be cool, and look smart, so you didn’t ask any questions. And now it’s too late. You’re stuck. Now you simply can’t ask … you’re too afraid. I get it. Apache Arrow is probably one of those things. It keeps popping […] The post What is Apache Arrow?

IT 130

More Trending

article thumbnail

Increase Your Odds Of Success For Analytics And AI Through More Effective Knowledge Management With AlignAI

Data Engineering Podcast

Summary Making effective use of data requires proper context around the information that is being used. As the size and complexity of your organization increases the difficulty of ensuring that everyone has the necessary knowledge about how to get their work done scales exponentially. Wikis and intranets are a common way to attempt to solve this problem, but they are frequently ineffective.

article thumbnail

More Data Science Cheatsheets

KDnuggets

It's time again to look at some data science cheatsheets. Here you can find a short selection of such resources which can cater to different existing levels of knowledge and breadth of topics of interest.

article thumbnail

Looking to the Future – How a Data Operating System Breathes Life Into Healthcare

The Modern Data Company

Looking to the Future – How a Data Operating System Breathes Life Into Healthcare Download (PDF) The post Looking to the Future – How a Data Operating System Breathes Life Into Healthcare appeared first on TheModernDataCompany.

article thumbnail

The Top Data Strategy Influencers and Content Creators on LinkedIn

Databand.ai

The Top Data Strategy Influencers and Content Creators on LinkedIn Eitan Chazbani 2022-12-29 14:08:41 What’s the latest in the data world? In a space that moves at a rapid-fire pace, keeping up with new trends and evolving best practices can be dizzying. But having the right network can make all the difference. Regularly following updates from leaders in the data strategy space can go a long way toward not only helping you stay up to date on the latest and greatest, but also allowing you to join

BI 52
article thumbnail

Prepare Now: 2025s Must-Know Trends For Product And Data Leaders

Speaker: Jay Allardyce, Deepak Vittal, and Terrence Sheflin

As we look ahead to 2025, business intelligence and data analytics are set to play pivotal roles in shaping success. Organizations are already starting to face a host of transformative trends as the year comes to a close, including the integration of AI in data analytics, an increased emphasis on real-time data insights, and the growing importance of user experience in BI solutions.

article thumbnail

Simple And Scalable Encryption Of Data In Use For Analytics And Machine Learning With Opaque Systems

Data Engineering Podcast

Summary Encryption and security are critical elements in data analytics and machine learning applications. We have well developed protocols and practices around data that is at rest and in motion, but security around data in use is still severely lacking. Recognizing this shortcoming and the capabilities that could be unlocked by a robust solution Rishabh Poddar helped to create Opaque Systems as an outgrowth of his PhD studies.

article thumbnail

Top 38 Python Libraries for Data Science, Data Visualization & Machine Learning

KDnuggets

This article compiles the 38 top Python libraries for data science, data visualization & machine learning, as best determined by KDnuggets staff.

article thumbnail

Building a Future in Banking and Capital Markets

The Modern Data Company

Banking and Capital Markets are undergoing a period of transformation. The global economic outlook is somewhat fragile, but banks are in an excellent position to survive and thrive as long as they have the right tools in place. According to Deloitte’s report 2023 Banking and Capital Markets Outlook , banks must find ways to adapt to global disruption and understand the changing needs of consumers to find success.

Banking 52
article thumbnail

Top 5 Data Engineering Deep Dives in 2022

Monte Carlo

No one wants to read marketing fluff, especially not data engineers. These builders and architects are prone to scoff at any article detailing concepts at a “high-level.” Everyone understands that data lineage and data pipeline monitoring are important, but the real question is, “how do you build it?” Caveat emptor, the following articles are for the technically inclined and definitely not for the faint of heart.

article thumbnail

How to Drive Cost Savings, Efficiency Gains, and Sustainability Wins with MES

Speaker: Nikhil Joshi, Founder & President of Snic Solutions

Is your manufacturing operation reaching its efficiency potential? A Manufacturing Execution System (MES) could be the game-changer, helping you reduce waste, cut costs, and lower your carbon footprint. Join Nikhil Joshi, Founder & President of Snic Solutions, in this value-packed webinar as he breaks down how MES can drive operational excellence and sustainability.

article thumbnail

An Exploration Of Tobias' Experience In Building A Data Lakehouse From Scratch

Data Engineering Podcast

Summary Five years of hosting the Data Engineering Podcast has provided Tobias Macey with a wealth of insight into the work of building and operating data systems at a variety of scales and for myriad purposes. In order to condense that acquired knowledge into a format that is useful to everyone Scott Hirleman turns the tables in this episode and asks Tobias about the tactical and strategic aspects of his experiences applying those lessons to the work of building a data platform from scratch.

Building 100
article thumbnail

Key Data Science, Machine Learning, AI and Analytics Developments of 2022

KDnuggets

It's the end of the year, and so it's time for KDnuggets to assemble a team of experts and get to the bottom of what the most important data science, machine learning, AI and analytics developments of 2022 were.

article thumbnail

The Terms and Conditions of a Data Contract are Data Tests

DataKitchen

The Terms and Conditions of a Data Contract are Automated Production Data Tests. A data contract is a formal agreement between two parties that defines the structure and format of data that will be exchanged between them. Data contracts are a new idea for data and analytic team development to ensure that data is transmitted accurately and consistently between different systems or teams.

article thumbnail

Our Top 5 Data Mesh Articles In 2022

Monte Carlo

Data mesh is a complex socio-technological data engineering concept, but it doesn’t change too much. The four principles are still the four principles, there are still three experience planes, and automation is still as vital as ever. This is a good thing! Data mesh is one of those rare transformative concepts that emerged relatively fully formed as a result of creator Zhamak Dhegani’s years of consulting experience captured in a comprehensive 384 page book.

Retail 52
article thumbnail

Improving the Accuracy of Generative AI Systems: A Structured Approach

Speaker: Anindo Banerjea, CTO at Civio & Tony Karrer, CTO at Aggregage

When developing a Gen AI application, one of the most significant challenges is improving accuracy. This can be especially difficult when working with a large data corpus, and as the complexity of the task increases. The number of use cases/corner cases that the system is expected to handle essentially explodes. 💥 Anindo Banerjea is here to showcase his significant experience building AI/ML SaaS applications as he walks us through the current problems his company, Civio, is solving.

article thumbnail

Snowflake: SSE File Encryption using AWS KMS

Cloudyard

Read Time: 3 Minute, 2 Second SSE File Encryption: During this post we will discuss an ERROR while executing the COPY command. Recently we got an issue while loading data from S3 bucket to Snowflake. According to the scenario, there were two files present in the bucket but surprisingly COPY command was failing to process one File. The command was reporting Access denied error for particular file.

AWS 52
article thumbnail

Data-Driven Holiday Cheer: How Santa is Using Analytics to Make the Season Bright

KDnuggets

Want to know how Santa might use data science to make his job easier? So did we, so we asked ChatGPT. Read on to find out what it said.

article thumbnail

How to Execute Linux Commands in Python?

Workfall

Reading Time: 8 minutes As of this writing, Linux has a global desktop market share of 2.77% ( A Report by Statcounter ), but it powers over 90% of all cloud infrastructure and hosting services. It is critical to be familiar with common Linux commands for this reason alone. According to a 2022 StackOverflow survey , Linux-based operating systems are more popular than macOS, demonstrating the appeal of using open-source software by professional developers, with an impressive 39.89% market share.

Python 52
article thumbnail

Holiday Downtime, Without Data Downtime

The Modern Data Company

Data center downtime can be costly. Gartner estimates that downtime can cost $5,600 per minute, extrapolating to well over $300K per hour. When your organization’s digital service is interrupted, it can impact employee productivity, company reputation, and customer loyalty. It can also result in the loss of business, data, and revenue. With the heart of the holiday season happening, we have tips on how to enjoy holiday downtime while avoiding the high costs of data center downtime.

article thumbnail

The Ultimate Guide To Data-Driven Construction: Optimize Projects, Reduce Risks, & Boost Innovation

Speaker: Donna Laquidara-Carr, PhD, LEED AP, Industry Insights Research Director at Dodge Construction Network

In today’s construction market, owners, construction managers, and contractors must navigate increasing challenges, from cost management to project delays. Fortunately, digital tools now offer valuable insights to help mitigate these risks. However, the sheer volume of tools and the complexity of leveraging their data effectively can be daunting. That’s where data-driven construction comes in.

article thumbnail

How Data Products Are Changing Market Economics

Acceldata

From a build perspective, data products ultimately translate into products that utilize data to improve services and overall functionality. And if we were to go by this definition, it becomes clear that no product in the world can truly survive unless they are a “data product”.

article thumbnail

The Zen of Python

KDnuggets

Python is one of the programming languages that are very versatile and relatively easy to learn. Hence it is the choice of many new programmers, regardless of what area of tech they are interested in. It is particularly popular in all data science branches.

Python 106
article thumbnail

How to Solve 4 Elasticsearch Performance Challenges at Scale

Rockset

Scaling Elasticsearch Elasticsearch is a NoSQL search and analytics engine that is easy to get started using for log analytics, text search, real-time analytics and more. That said, under the hood Elasticsearch is a complex, distributed system with many levers to pull to achieve optimal performance. In this blog, we walk through solutions to common Elasticsearch performance challenges at scale including slow indexing, search speed, shard and index sizing, and multi-tenancy.

article thumbnail

Our Top 5 Most Popular Data Engineering Articles In 2022

Monte Carlo

The Pareto Principle , which holds 80% of the results will derive from 20% of the cases, is tough to escape. It definitely holds true for our Data Downtime blog with these five articles driving a majority of our traffic in 2022. There are a few characteristics that separate these articles from the chaff, namely: They were among the first to describe or even define a nascent concept.

article thumbnail

Business Intelligence 101: How To Make The Best Solution Decision For Your Organization

Speaker: Evelyn Chou

Choosing the right business intelligence (BI) platform can feel like navigating a maze of features, promises, and technical jargon. With so many options available, how can you ensure you’re making the right decision for your organization’s unique needs? 🤔 This webinar brings together expert insights to break down the complexities of BI solution vetting.

article thumbnail

Best of 2022: Round Up

Precisely

As 2022 wraps up, we would like to recap our top posts of the year in Data Integrity, Data Integration, Data Quality, Data Governance, Location Intelligence, SAP Automation, and how data affects specific industries. Let’s take a look! Best of Data Integrity Data integrity empowers your businesses to make fast, confident decisions based on trusted data that has maximum accuracy, consistency, and context.

article thumbnail

5 Tasks To Automate With Python

KDnuggets

Here are 5 tasks you can automate with Python, and how to do it.

Python 160
article thumbnail

Data Engineering Weekly in Year 2022

Data Engineering Weekly

The holidays bring joy and memories. It is always a joyful memory for me every week when I pen down (or key down 🤷🏽‍♂️) every edition of Data Engineering Weekly. I want to take a holiday break for this week's edition, and instead, I want to reflect on our journey in 2022. A Growth To Remember 2022 has been a remarkable year in terms of subscriber growth.

article thumbnail

Barr Moses: My Top 5 Articles of 2022

Monte Carlo

I don’t see myself as a writer or blogger. In fact, the first blog post I published on Medium sat as a draft for months. ( Data downtime , anyone?) Prior to launching Monte Carlo, I interviewed hundreds of data leaders. I gained so much insight into their hopes, dreams, and fears that the impulse to share finally exceeded the anxiety of publishing. And there was no turning back.

article thumbnail

Driving Responsible Innovation: How to Navigate AI Governance & Data Privacy

Speaker: Aindra Misra, Senior Manager, Product Management (Data, ML, and Cloud Infrastructure) at BILL

Join us for an insightful webinar that explores the critical intersection of data privacy and AI governance. In today’s rapidly evolving tech landscape, building robust governance frameworks is essential to fostering innovation while staying compliant with regulations. Our expert speaker, Aindra Misra, will guide you through best practices for ensuring data protection while leveraging AI capabilities.