Showing posts with label storage. Show all posts
Showing posts with label storage. Show all posts

Thursday, June 26, 2014

AWS encrypting data at rest

Here is good white paper on encrypting data at rest on AWS:
http://media.amazonwebservices.com/AWS_Securing_Data_at_Rest_with_Encryption.pdf

Amazon now offers Amazon EBS native encryption: http://aws.amazon.com/about-aws/whats-new/2014/05/21/Amazon-EBS-encryption-now-available/
S3 has SSE encryption, client side encryption and SSE with key managed by you: http://aws.amazon.com/blogs/aws/s3-encryption-with-your-keys/
All data in Glacier and Redshift is automatically encrypted.

Thursday, April 24, 2014

AWS Import/Export file limits

AWS import/Export can handle volumes of larger then 1TB to be stored on Amazon EBS volumes. However, there is a intermediate step using S3.  If your storage device’s capacity is less than or equal to the maximum Amazon EBS volume size of 1TB, its contents will be loaded directly into an Amazon EBS snapshot. So, in theory no size limit. AWS does not mount the file system on your storage device, nor is a file system required to be present. AWS Import/Export performs a block for block copy from your device to an Amazon EBS Snapshot. If your storage device’s capacity exceeds 1TB, a device image will be stored within your specified Amazon S3 log bucket. You can then create a RAID of EBS volumes using software such as Logical Volume Manager, and copy the image from Amazon S3 to this new volume

Sunday, March 30, 2014

Redshift cluster sizes

100 nodes maximum for each configuration.
Dense Storage (DW1) nodes are available in two sizes.   These are HDD backed instances.
A. The Extra Large has three HDDs with a total of 2TB of magnetic storage. Maximum of 200 TB of storage.
B. The Eight Extra Large has 24 HDDs with a total of 16TB of magnetic storage. Maximum of 1.6 Pedabyte of storage. 
Dense Compute (DW2) nodes are also available in two sizes.  These are SSD back instances.
A. The Large has 160GB of SSD storage per EC2 instance with a maximum is 1.6 TB
B The Eight Extra Large is sixteen times bigger (then the dense compute large) with 2.56TB of SSD storage on the EC2 instance for a maximum of 256TB of SSD storage.

Friday, March 28, 2014

NAS and SAN on AWS


Network Attached Storage (NAS) file level storage. Storage Area Network (SAN) block level storage.  NAS lends itself to quickly provisioning an exported NFS file system or CIFS share to a user backed by a thinly provisioned physical file system. SANs typically use iSCSI.  AWS does not offer native shared / clustered disk that on premise NAS and SAN offerings. Here are some options for running NAS and SAN environments in AWS:

Here is a good session from reInvent in 2013 (keep in mind this information has changed since over a year old) : http://www.slideshare.net/AmazonWebServices/nfs-and-cifs-options-for-aws-stg401-aws-reinvent-2013

Of course, S3 can always be used but this will require changes to the application or use an emulation software (s3fs) to make S3 look like block/file level storage. 

Thursday, January 2, 2014

Accenture white paper : Disaster Recovery with Amazon Web Services


Disaster Recovery with Amazon Web Services:
A Technical Guide
The paper covers
1. Define the challenges that enterprises face in adopting public cloud solutions for disaster recovery.
2. Describe the value that large enterprises can gain by adopting cloud-based DR with services such as Amazon Web Services (AWS) for disaster recovery.
3. Provide recommended disaster recovery architecture patterns.

http://www.accenture.com/microsite/reinvent-2013/Documents/Accentue-Smart-Disaster-Recovery-with-Amazon-Web-Services.pdf

Friday, December 20, 2013

Oracle Database on SSD

 New options for running the Oracle Database on EC2 using SSD storage for your database. The I2 instances feature the latest generation of Intel Ivy Bridge processors - each virtual CPU (vCPU) is a hardware hyperthread from an Intel Xeon E5-2670 v2 (Ivy Bridge) processor. I2 instances are available in four sizes as listed in the table below.
Instance TypevCPUsECU RatingMemory (GiB)Instance Storage SSD (GB)Note          
i2.xlarge41430.51 x 800800 GB of SSD
i2.2xlarge827612 x 8001.6 TB of SSD
i2.4xlarge16531224 x 8003.2 TB of SSD. hi.4xlarge with 2 TB  of SSD was previously largest SSD instance type
i2.8xlarge321042448 x 8006.4 TB of SSD


The i2.8xlarge instance size is capable of providing over 365,000 4 kilobyte (KB) random read IOPS and over 315,000 4 KB random write IOPS.   Great for Oracle databases that have high IO requirements.  This document has more on configuring Oracle Databases on EC2 and using SSD based instances : http://cloudconclave.blogspot.com/2013/11/aws-database-reference-implementation.html
 

Tuesday, December 3, 2013

On premise mapped to AWS

A question that often comes up when companies are migrating Oracle workloads to AWS is: "How do my on premise IT architecture components map to AWS?".  Here are some of the most common components mapping from on premise to AWS:



Monday, December 2, 2013

Storage tiering on AWS

Here is the replay of the session I just presented with an AWS partner (App Associates) and customer (Riso):http://cloudconclave.blogspot.com/2013/11/aws-storage-tiering-for-enterprise.html

I opened the session by 'testing' the audiences understanding of AWS storage tiers:
1. Assuming a 16K block size, what storage option provides average throughput of 1 to 2 MBPS ?
2. What storage option has single threaded through put of around 17 MBPS ?
3. Assuming a 16K block size, what EBS volume provides an average of 16-20 MBPS through put ?
4. Once again assuming 16K block size and also assuming 4 1K PIOPS volumes, for what EC2 instances will you start to see network saturation ?
5. What storage option produces approximately 100-145 MBPS read and write through put?
6. For which storage option is it possible to transfer approximately 3 TB a day a day over a WAN ?




















Answers:
1. Standard IOPS
2. S3
3. 1K PIOPS
4. Any instance with .5 Gpbs network connection. For example, m2.2xlarge
5. Hi1.4xlarge (high IO) ephemeral storage
6. AWS Storage gateway


Oracle Database size and network throughput : IOPS and network throughput

When configuring an Oracle Database on AWS EC2, you need to consider both storage (EBS) IOPS and network throughput.   With PIOPS, you can achieve up to 4K IOPS per EBS volumes.  However, the EC2 instances (assuming EBS optimized or 10 Gbps cluster compute) can be a potential bottle neck.  For example, the m2.2xlarge instance type has a maximum throughput of 0.5Gb.  Which means this instance type is limited to approximately 3750 * 16 KiB IOPS. Therefore, one 4K PIOPS volume would start to saturate the network.   Take into account the IOPS and instance network throughput when designing your Oracle Database on AWS EC2.

Data stores compatible with Amazon EMR

There are a number of different file systems that can be used

1. Hadoop Distributed File System (HDFS) : EC2 local/ephemeral disk is where HDFS  resides.  The obvious disadvantage is that it’s ephemeral storage which is reclaimed when the cluster ends. It can be used for caching the results produced by intermediate job-flow steps during a large EMR job.
2. Local (ephemeral) EC2 disk :  Each EMR node comes with local disk.  This disk works well for temporary storage of data that is continually changing, such as buffers, caches, scratch data, and other temporary content.
3. S3 native : Used for input (data set to be reduced) and output/results.
4. S3 block : Stay away from as not as performant as the other options.
5. HBase : HBase is an open source, non-relational, distributed database that runs on top of HDFS.  HBase works with Hadoop/EMR, sharing its file system and serving as a direct input and output to EMR jobs. HBase also integrates with Apache Hive, enabling SQL-like queries over HBase tables, joins with Hive-based tables, and support for Java Database Connectivity (JDBC).

More information here:
http://docs.aws.amazon.com/ElasticMapReduce/latest/DeveloperGuide/emr-plan-file-systems.html



Monday, November 18, 2013

AWS Database reference implementation

This reference implementation provides the architecture and associated CloudFormation templates for a standard, enterprise class, large enterprise class and high performance Oracle 11g configuration on AWS EC2:
http://media.amazonwebservices.com/AWS_RDBMS_Oracle_11g_on_EC2_Reference_Architecture.pdf

Monday, September 30, 2013

AWS EBS PIOPS : block size and IOPS


Having spent more time in the database world than in the web development world, I am accustomed to measuring (database) performance/through put in terms of IOPS or TPS.  The web/video/image world like to use MB/sec.  Why I am saying this? Because it relates to the conversation about getting a certain level of PIOPS (based upon a 16 KB block) on AWS EBS and how this effects MB/sec.  MB/sec, I am beginning to understand, and maybe move to the 'dark side', is the ultimate measure of disk 'performance'.  

Example: A 2000 Provisioned IOPS volume can handle:
•2000 16KB read/write per second, or 1000 32KB read/write per second, or 500 64KB read/write per second 
•You will get consistent 32 MB/sec throughput (with 16KB or higher IOs)
•Perform an index creation action and sends I/O of 32K, IOPS becomes 1000, you still get 32MB/sec throughput
•On best effort, you may get up to 40 MB/sec throughput 

So, you may be better off using a 64 KB block size but your PIOPS will show up as lower but your MB/sec could be better.

Friday, August 30, 2013

AWS re:Invent enterprise sessions


DAT202 - Using Amazon RDS to Power Enterprise Applications Amazon Relational Database Service (Amazon RDS) makes it cheap and easy to deploy, manage, and scale relational databases using a familiar MySQL, Oracle or MS SQL server database engine. Amazon RDS can be an excellent choice for running many large, off-the-shelf enterprise applications from companies like JD Edwards, Oracle, PeopleSoft, and Siebel. Sign up for this session to learn how to best leverage Amazon RDS for use with enterprise applications and learn about best practices and data migration strategies.


DAT401 - Advanced Data Migration Techniques for Amazon RDS Migrating data from the existing environments to AWS is a key part of the overall migration to Amazon RDS for most customers. Moving data into Amazon Relational Database Service (Amazon RDS) from existing production systems in a reliable, synchronized manner with minimum downtime requires careful planning and the use appropriate tools and technologies. Because each migration scenario is different in terms of source and target systems, tools, and data sizes, you'll need to customize your data migration strategy to achieve the best outcome. In this session we will do a deep dive into various methods, tools and technologies that can be put to use for a successful and timely data migration to Amazon RDS.


STG301 - AWS Storage Tiers for Enterprise Workloads - Best Practices Enterprise environments utilize many different kinds of storage technologies from different vendors to fulfill various needs in their IT landscape. These are often very expensive and procurement cycles quite lengthy. They also need specialized expertise in each vendor's storage technologies to configure them and integrate them into the ecosystem, again resulting in prolonged project cycles and added cost. AWS provides end-to-end storage solutions that will fulfill all these needs of Enterprise Environments that are easily manageable, extremely cost effective, fully integrated and totally on demand. These storage technologies include Elastic Block Store (EBS) for instance attached block storage, Amazon Simple Storage Service (Amazon S3) for object (file) storage and Amazon Glacier for archival. An enterprise database environment is an excellent example of a system that could use all these storage technologies to implement an end-to-end solution using striped PIOPS volumes for data files, Standard EBS volumes for log files, S3 for database backup using Oracle Secure Backup and Glacier for long-time archival from S3 based on time lapse rules. In this session, we will explore the best practices for utilizing AWS storage tiers for enterprise workloads

STG305 - Disaster Recovery Site on AWS - Minimal Cost Maximum Efficiency Implementation of  disaster recovery (DR) site is crucial for business continuity of any enterprise. Due to the fundamental nature of features like elasticity, scalability and geographic distribution, DR implementation on AWS can be done at 10-50% of the conventional cost. In this session, we will do a deep dive into proven DR architectures on AWS and best practices, tools and techniques get the most out of them.


STG303 - Running Microsoft and Oracle Stacks on Elastic Block Store Run your Enterprise applications on Elastic Block Store (EBS). This session will discuss how you can leverage the block storage platform (EBS) as you move your Microsoft (SQL Server, Exchange, SharePoint) and Oracle (Databases, E-business Suite, Business Intelligence) workloads onto Amazon Web Services (AWS). The session will cover high availability, performance, and backup/restore best practices

ENT303 - Migrating Enterprise Applications to AWS - Best Practices, Tools and Techniques In this session we will discuss strategies, tools and techniques for migrating enterprise software systems to AWS. We'll consider applications like Oracle eBusiness Suite, SAP, PeopleSoft, JD Edwards and Siebel. These applications are complex by themselves; they are frequently customized; they have many touch points on other systems in the enterprise; and they often have large associated databases. Nevertheless, running enterprise applications in the cloud affords powerful benefits. We'll identify success factors and best practices.