Saturday, May 20, 2017

Hadoop Hive, local setup

I have spent many hours to get hive run locally in MacOS but couldn’t make it. Last time I get to the very end of this tutorial except the last step. This time I proceed a little further but bugs keep poping up:
$ hdfs dfs -mkdir /user
Cannot create directory /user. Name node is in safe mode.
$ hdfs dfsadmin -safemode leave
Safe mode is OFF
$ hdfs dfsadmin -safemode get
Safe mode is ON
Anyway, I try to record every step of my journey.
According to Quora, the minimum requirement for a local machine is 500 GB. This may be reason I failed.

Download tar files from respective official sites:
  1. oracle Java SE
  2. hadoop: http://hadoop.apache.org/releases.html
  3. hive
  4. derby

bash command line refresh

export varname=value  # export a variable to environment
env  # disply all environment variables, note that different shells have different default env variables
cat .bash_profile  # see a file in command window
less .bash_profile  # another way to see, less overwhelming
echo $varname # display variable value, note the dollar sign
eval $fun  # evaluate function
history  # display command history
hash     # display command history and path
pwd  # equal to echo $PWD which is a buit-in variable
let arg1=2  # define variable value, space is forbidden
let arg2=$arg1**3
echo $arg2
printf "result=%d\n" $arg2
Most compiler/commands are stored at /usr/local/bin .

path setup

# setup environment for hadoop
export HADOOP_HOME=/usr/local/hadoop-2.8.0    
export HADOOP_MAPRED_HOME=$HADOOP_HOME
export HADOOP_COMMON_HOME=$HADOOP_HOME
export HADOOP_HDFS_HOME=$HADOOP_HOME
export YARN_HOME=$HADOOP_HOME
export HADOOP_COMMON_LIB_NATIVE_DIR=$HADOOP_HOME/lib/native 
export PATH=$PATH:$HADOOP_HOME/sbin:$HADOOP_HOME/bin

# setup environment for hive
export HIVE_HOME=/usr/local/apache-hive-2.1.1-bin 
export PATH=$PATH:$HIVE_HOME/bin
export CLASSPATH=$CLASSPATH:/usr/local/hadoop-2.8.0/lib/*:.
export CLASSPATH=$CLASSPATH:/usr/local/hive-2.1.1/lib*:.

# setup environment for Derby
export DERBY_HOME=/usr/local/db-derby-10.13.1.1-bin
export PATH=$PATH:$DERBY_HOME/bin:$HIVE_HOME/bin
export CLASSPATH=$CLASSPATH:$DERBY_HOME/lib/derby.jar:$DERBY_HOME/lib/derbytools.jar

hadoop initialize and commands

cd /usr/local/hadoop-2.8.0/
hdfs namenode -format
sbin/start-dfs.sh   # start Hadoop file system
# open http://localhost:50070/  
sbin/start-yarn.sh
# open http://localhost:8088/  

hadoop fs -mkdir /tmp 
hadoop fs -mkdir -p ~/hive/warehouse #also make pararent dir 
hadoop fs -chmod 777 /user  # change permission of file or folder
hdfs dfs -mkdir /user/hadoop  # make folder
hdfs dfs -put a.csv /user/hadoop/a.csv # move from local to HDFS
hdfs dfs -ls /user/hadoop  # list content of a folder
hdfs dfs -du  /user/hadoop/  # display utilization (size)
hdfs dfs -get /user/hadoop/ /home/ # get from HDFS to local
hdfs dfs -cp /user/hadoop/folderA /user/hadoop/folderB # copy
hdfs fs -rm -r <directory>  # remove

Hive metastore_db initialize

schematool -initSchema -dbType derby # may fail
mv metastore_db metastore_db.tmp #
schematool -initSchema -dbType derby #rerun
hive
show tables;
create table myGod (name string);
hive metastore configuration
add follows to hive-site.xml
<property> 
<name>system:java.io.tmpdir</name> 
<value>/usr/local/apache-hive-2.1.1-bin /iotmp</value> 
<description/> 
</property>
Hive use Derty database as default. You may change it to mySQL database by following the above link.
​

Sunday, May 14, 2017

Hadoop Hive setup by Cloudera quickstart


For beginners of Hadoop and Hive, a good starting point is to use Cloudera quickstart. Because the tricky configuration and overwhelming warning may scare off beginners. Steps:
  1. Download virtual box.
  2. Download cloudera quickstart vm at https://www.cloudera.com/downloads/
  3. use import .ovf file for setup. I was stupid to try the manual setup.
  4. start the vm.
  5. In the pop-up firefox browser, go through quickstart.cloudera tutorial to get yourself familiar with popular Hadoop framework/tools/concepts: Hue, Hive, file browser, sqoop, impala, parquet
Following is my learning notes for the tutorial

1, Ingest and Query Relational Data

Use Apache Sqoop to load relational data from MySQL into HDFS. With a few additional parameters, the relational data can be ready to be queried by Impala with Hadoop optimized file format Apache Avro.
sqoop import-all-tables \    -m 1 \    --connect jdbc:mysql://quickstart:3306/retail_db \    
--username=retail_dba \    --password=cloudera \    
--compression-codec=snappy \    
--as-parquetfile \    
--warehouse-dir=/user/hive/warehouse \    
--hive-import
It is launching MapReduce jobs to pull the data from our MySQL database and write the data to HDFS, distributed across the cluster in Apache Parquet format. Parquet is a format designed for analytical applications on Hadoop. Instead of grouping your data into rows like typical data formats, it groups your data into columns. This is ideal for many analytical queries where instead of retrieving data from specific records.
Hue provides a web-based interface for many of the tools in CDH with address: quickstart.cloudera:8888. In the QuickStart VM, the administrator username for Hue is ‘cloudera’ and the password is ‘cloudera’.
we told Sqoop to import the data into Hive but used Impala to query the data. This is because Hive and Impala can share both data files and the table metadata. Hive works by compiling SQL queries into MapReduce jobs, which makes it very flexible, whereas Impala executes queries itself and is built from the ground up to be as fast as possible, which makes it better for interactive analysis. We’ll use Hive later for an ETL (extract-transform-load) workload.
Simply put, Hive aims for compatibility and Impala aims for speed.

2, Correlate Structured Data with Unstructured Data

use hive to parse the unstructured log data and use impala to query.First create a table by parsing data using regular expression
CREATE EXTERNAL TABLE intermediate_access_logs (    ip STRING,    date STRING,    method STRING,    url STRING,    http_version STRING,    code1 STRING,    code2 STRING,    dash STRING,    user_agent STRING) ROW FORMAT SERDE 'org.apache.hadoop.hive.contrib.serde2.RegexSerDe'WITH SERDEPROPERTIES ('input.regex' = '([^ ]*) - - \\[([^\\]]*)\\] "([^\ ]*) ([^\ ]*) ([^\ ]*)" (\\d*) (\\d*) "([^"]*)" "([^"]*)"',    'output.format.string' = "%1$$s %2$$s %3$$s %4$$s %5$$s %6$$s %7$$s %8$$s %9$$s") LOCATION '/user/hive/warehouse/original_access_logs';
create another table, and then
INSERT OVERWRITE TABLE tokenized_access_logs SELECT * FROM intermediate_access_logs;
This does MapReduce job. You will get a table with 160 k records.Once it is ready, we can use impala to do the query:
select count(*),url from tokenized_access_logswhere url like '%\/product\/%'group by url order by count(*) desc;
Then analyze why some products are viewed most but don’t have a good sale.

3. Relationship strength analytics using Spark

The tool in CDH best suited for quick analytics on object relationships is Apache Spark.

4. Explore Log Events Interactively

learned how to use Cloudera Search to allow exploration of data in real time, using Flume and Solr and Morphlines

5. Hue Dashboard for data visualization

The plot is relatively simple. Maybe because it’s a free version.

short history of Hadoop

  • In March 2013 Intel invested $740 million in Cloudera for an 18% investment. Intel shared its roadmap with Cloudera so Cloudera could develop the software to maximize the chip performance. With cash piled from other investors by selling equity, Cloudera is able to acquire other company when necessary.
  • super evangelical to do a technology education project in order to win early customers.
  • nobody buys a database because it is easy to manage. people buy applications. They have important business or mission or operational problem to solve.So they buy the software that allows a non-programmer to do the job. That requires applications tools systems integrated together stack on top of the database.
  • Since the mid-1980s that dynamic is very much alive in this ecosystem. People bring existing skills to this new platform. BI, data analytics report, machine learning that you can never do that before but you can do it now.
  • banks, hospital, retail stores take security very seriously. If you break into a yahoo cluster steal a new story, who cares? Break into a bank’s big data platform, steal some transaction data that’s a big issue.
  • the question is ill-formed. Open source is a distribution model, license model, a development model but not a business model. You can’t build the long-term sustainable pure open-source business. Open source project often gets acquired by big company which has other revenue stream.
  • my competitors will tell you that I’m a furious guy that’s trying to lock our customers in because of my evil desires on their wallets. I suspect so. I’ve heard rumors. I am trying to lock IBM out. We will always have proprietary IP at Cloudera.
  • Who can afford to hire the thousands of smart people around the world working on this thing(open source)? No single company can compete with that. And we get the benefit.
  • plan to IPO. Cloudera was just traded at NYSE as CLDR on 2017.4.28, with IPO price at $18.

A New Generation Of Data Scientists

  • Most of the time I don’t do the fancy data visualization as in Sci-fi. I do data cleansing, prepare for the dataset.
  • big data economics: no individual record is particularly valuable. Having every record is incredibly valuable.
  • google file system: 4 kB per block. HDFS: 64/256 MB per block.
  • here’s a bunch of data and find me some insights. This is the worst thing can happen to me. I would say to the business person: tell me the problem you have. In the process of solving problem, we will discover the insights. Insights don’t come from vacuum. Insights come from interesting meaningful problems.
  • Have a data science team. Nobody is good at all the skills. You have to know every part and talk to everybody in the team.
  • we have to have skills not only analyzing data, but also deploy model in production system. Being data scientist involves some of the skills in DevOps.
  • Don’t solve the problem once. Solve it zero or solve it with multiple models. So choose the good problem first.
  • It’s never the case that we are trying to optimize a single thing in a data science problem. Like in an Ad prediction model, it’s not only use machine learning model to predict the click, but how to increase the revenue.
  • be self. iterate until awesome.
  • Hadoop developer training, hive and pig training, intro to data science (end to end problem-solving). recommendation system is in every field.

some terminology

ETL(Extract, Transform, Load): a repeatable programmed data movement
Extract: get data from source, is the most resource intensive
Transform: filter/map/enrich/combine/validate/sort, most difficult
Load: store data in a data warehouse or data mart.
Apache Spark was a cluster-computing framework, released in 2014.5.30 to address the limitation in the MapReduce cluster computing paradigm, which is a linear data flow structure. Spark provides a data structure called resilient distributed dataset (RDD), which facilliates the iterative algorithms of data access and data analysis.
Hive is not designed for online transaction processing. It is best used for traditional data warehousing tasks.
​

Friday, May 12, 2017

Intro to Fortran


Due to a technical interview at Intel, I have to pick up this legacy language.
Fortran, derived from Formula Translation, is a general-purpose, imperative programming language that is especially suited to numeric computation and scientific computing. It was originally developed by IBM in the 1950s. It has still been used in computationally intensive areas due to its fast speed and existing software/packages.
.f file is used in early Fortran program written in a fixed-column format to reflect the 80-column punched-card practice. So it has very weird grammar:
  • in each line, first 1-5 are label fields, it can be c (comment) or number (notation for the code block)
  • 6th column. If it is something other than 0, it means the code continues from the previous line
  • 7~72. Real, independent codes
  • 73-80 are ignored because the IBM 704 card reader only had 72 columns
Smiley face
After Fortran 90, the Free Format is used and the file extension is .f90 In this format, the comment is signaled by ! and each line can be 132 symbols without the need of first 5 empty columns. The between line continuation is signaled by & at the end of the previous line as well as the head of next line.
Most commonly used versions today are: Fortran 77, Fortran 90, and Fortran 95. Newer versions such as Fortran 2008 only adds minor revision.

setup

There are several ways to install Fortran compiler/IDE.

1. GNU compiler

brew install gcc
🍺 /usr/local/Cellar/gcc/7.1.0: 1,485 files, 289.6MB
This GNU version of compiler bundles fortran and c together.

2. intel compiler

Intel® Parallel Studio XE Composer Edition for Fortran macOS*
Serial number : 26BK-MCT25TSK
expire 2018-5-12
But it turns out to be a compiler, which must be used along with Microsoft visual studio or Max Xcode.

3. eclipse Photran

run compiler

gfortran xx.f      # default output file is named a.out
gfortran xx.f -o xx  # customerize file name
./a.out   # execute file

tutorial

stanford

From time to time, so-called experts predict that Fortran will rapidly fade in popularity and soon become extinct. These predictions have always failed. Fortran is the most enduring computer programming language in history. One of the main reasons Fortran has survived and will survive is software inertia. Once a company has spent many man-years and perhaps millions of dollars on a software product, it is unlikely to try to translate the software to a different language. Reliable software translation is a very difficult task.
Use Fortrain 77 compiler on a Unix workstation.
Install libraries? Libraries have file names starting with lib and ending in .a. Some libraries have already been installed by your system administrator, usually in the directories /usr/lib and /usr/local/lib. For example, the BLAS library MAY be stored in the file /usr/local/lib/libblas.a. You use the -l option to link it together with your main program, e.g.
      f77 main.f -lblas

tutorials point

This website provides an online fortran 95 environment.
program title
implicit none  ! let compiler check all variables
real :: a, b, result ! declare variable type
a = 12.0
b = 15.0
result = a + b
print *, "the total is", result  ! * means format
write(*,*) reult ! similar to print, more variety
end program title  ! finish program
fortran is case insensitive.
variable type:
integer a
a = 1
real b
b = 1.0
real(kind=8) c ! declare bite size
c = 1e9
double precision cc ! double precision
cc = 1.578d10
complex d 
d = (3.2,2.5)  ! set complex value
character e ! declare one lette
character(len=10) f ! declear string size
f = "Hello"
logical h
h = .true.  ! note the weird dots
real, parameter :: pi = 3.14159 ! declare tyes and set initial value, using two colons
integer i, j
equivalence (i,j)  ! using the same meomory
mod (b,c)  ! equivalent to % in python
customized type
type :: person  ! begin to define a type person
    character(len=30) :: name
    interger :: age
end type person  ! finish defining the type
type(person) :: a  ! declare a person type variable
write(*,*) "name:" ! prompt user to input name
read(*,*) a%name  ! read user input into name
logical control
if (a>b) then
    print *, "a is larger than b"
else if (a== b) then
    print *, "a is equal to b"
else
    print *, "a is not larger than b"
end if
I happen to have a book “Fortran 95 程序设计” (彭国伦) which I bought 5 years ago but never get a chance to read it until now. It turns out to be extremely good. It not only introduces FORTRAN 95, but also mentions between its improvement over Fortran 77, and how some old-fashion styles such as goto should be discarded. It really helps me to understand some legacy codes within a couple of hours. I remembered how the old-fashion formatted Fortran codes scared me off when I first encounter Fortran. I master it now.
​

Thursday, May 11, 2017

Linux Command Line Basics


For a Mac user who is familiar with bash shell, you already know a lot about the command line and how to install software in the black. What adds to your knowledge reserve is the server-side operation, which is becoming more common if handling bigger data.
Host OS: an operating system that’s installed directly on your physical computer.
Guest OS: OS installed indirectly, using a virtual machine (VM) software such as VirtualBox. VMs isolate programming projects from everything else without disrupting their day-to-day environment.
Install VirtualBox and Vagrant. My previous post has some more details when VM is used for psql.
Download the vagrantfile, cd to the target folder.
vagrant up   # download the 1.67 GB .vmdk file
vagrant ssh  # enter ubuntu 14, use 25% memory
ls
pwd
cd
curl http://udacity.github.io/ud595-shell/stuff.zip -o things.zip
sudo apt-get install cowsay
cowsay good morning
cowsay -e ^^ good morning
cowsay -f tux good morning
man conwsay
q  # exit manual
apropos working directory
bc # simple calculator
default shell on most Linux and Mac is GNU Bash.
bash --version
hostname
host udacity.com
date
history
rm xx.txt  # equal to os.remove("xx.txt")
uptime
(ctrl+r)  # search previous command
unzip things.zip
cat bivalves.txt  # read short file
less xx.txt  # read long file
wc bivalves.txt   # word count: lines,words,bytes
diff file1, file2  # show difference
nano xx.txt  # edit file. ctrl+key is shortcut
ping 8.8.8.8 # connect to another computer
  • use quote”” or backslash\ if the file name has special symbols or space
  • / is root, means absolute path
  • ../ is one-level up parent of current work directory
  • . path from root to current work directory
  • ~ / home directory
  • otherwise is a relative path, which is more convenient
package source list
cat /etc/apt/sources.list
sudo apt-get update   # update info
sudo apt-get upgrade  # upgrade software
sudo apt-get install finger
cat /etc/passwd  # record of all users !
sudo adduser student
ssh student@127.0.0.1 -p 2222
search package at http://packages.ubuntu.com
Linux distributions:
  • Red Hat: Enterprise level
  • Ubuntu: free and ease of use
  • Linux Mint Desktop users with proprietary media support
  • CoreOS: clustered, containerized deployment of apps.
​