AWS Big Data Blog

Category: Intermediate (200)

How Moeve standardized dbt runs across data lakes with Amazon Athena

Moeve standardized how it runs dbt across multiple data lakes by building a centralized, serverless launcher on Amazon Athena, AWS Step Functions, AWS Fargate, Amazon DynamoDB, and Amazon EventBridge, cutting new-project onboarding from days to about 15 minutes while keeping compute close to the data and orchestration loosely coupled.

Connect Amazon SageMaker Unified Studio to Microsoft Power BI - Part 1: IAM Identity Center (IDC)-based domains

Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 1: IAM Identity Center (IDC)-based domains

Connect Microsoft Power BI directly to governed data in Amazon SageMaker Unified Studio using new authentication modes in the Amazon Athena ODBC driver, with no third-party ODBC-JDBC bridge. Part 1 covers IAM Identity Center (IDC)-based domains with both DSN-based and DSN-less connection methods.

Connect Amazon SageMaker Unified Studio to Microsoft Power BI - Part 2: IAM-based domains

Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 2: IAM-based domains

Connect Microsoft Power BI directly to governed data in Amazon SageMaker Unified Studio using the Amazon Athena ODBC driver. Part 2 covers IAM-based domains with SageMakerIam authentication, including AWS IAM Identity Center administrator setup, for both DSN-based and DSN-less connection methods.

Deliver real-time data to streaming tables for Apache Iceberg with Amazon Kinesis Data Streams

Deliver real-time data to streaming tables for Apache Iceberg with Amazon Kinesis Data Streams

Amazon Kinesis Data Streams now supports streaming tables, a fully managed capability that continuously delivers your streaming data as queryable Apache Iceberg tables on Amazon S3 Tables. Streaming tables reduce data delivery costs to S3 Tables by up to 50% compared to self-managed alternatives and reduce downstream query costs by up to 30% through intelligent inline compaction that eliminates the small file problem. You need no custom applications, no self-managed compute, and no operational overhead.

PythonOperator and BashOperator now available on Amazon Managed Workflows for Apache Airflow (Amazon MWAA) Serverless

PythonOperator and BashOperator Now Available on Amazon Managed Workflows for Apache Airflow (Amazon MWAA) Serverless

You can now use PythonOperator and BashOperator to run custom Python functions and shell scripts directly in the Amazon MWAA Serverless runtime, without provisioning additional infrastructure. This post walks through building a serverless pipeline that converts CSV files to JSON using a PythonOperator and verifies the output with a BashOperator.

GPU-accelerated Apache Spark with Amazon EMR and NVIDIA RTX PRO 4500 on Amazon EC2 G7 instances runs up to 3.7x faster

GPU-accelerated Apache Spark with Amazon EMR and NVIDIA RTX PRO 4500 on Amazon EC2 G7 instances runs up to 3.7x faster

Amazon EMR on EKS now runs Apache Spark up to 3.7x faster on Amazon EC2 G7 instances with NVIDIA RTX PRO 4500 Blackwell GPUs than on comparable CPU instances, with no changes to existing Spark code. See the TPC-DS benchmark results, the cost comparison, and how to get started.

Long-term system tables retention in Amazon Redshift with Amazon S3 Tables

Long-term system tables retention in Amazon Redshift with Amazon S3 Tables

Amazon Redshift system table integration with Amazon S3 Tables automatically delivers your system table logs to Amazon S3 Tables in Apache Iceberg format. You can retain this data well beyond the 7-day limit for compliance, auditing, and cross-warehouse observability, without custom ETL pipelines or cluster resource consumption.

Track SageMaker Unified Studio project costs with custom tags and AWS CUR

Track SageMaker Unified Studio project costs with custom tags and AWS CUR

Learn how to track Amazon SageMaker Unified Studio project costs by custom tags. This serverless solution enriches AWS Cost and Usage Report (CUR) data with custom project tags and visualizes cost by CostCenter, Team, or Environment in an Amazon Quick Sight dashboard.

AI-powered cost optimization agent for Amazon Kinesis Data Streams

AI-powered cost optimization agent for Amazon Kinesis Data Streams

Learn how to deploy an open-source, AI-powered agent built on Amazon Bedrock that automatically analyzes every Amazon Kinesis Data Streams stream in your account, compares costs across the three capacity modes, and recommends the optimal mode to help you save over 60% on streaming costs on a schedule you choose.