CIS 612

Big Data and Parallel Distributed Data Processing Systems (3-0-3)

Course Content

  • Class Announcement and Post
  • Class Syllabus
  • Group Project and Presentation
  • Lab Assignments
  • Class Lecture Notes

  • Class_Notes: Introduction to Big Data Technologies and Big Data Processing Systems on Cloud
  • Class_Notes: Big Data and Processing Techniques as Semi-Structured Models
  • Class_Notes: Big Data and Processing Techniques as Unstructured Texts
  • Class_Notes: Big Data and NoSQL Big Data Systems:MongoDB
  • Class_Notes: Unstructured Data II: Knowledge Representation and Information Extraction Methods from Unstructured Data
  • Class_Notes: NLP with Machine Learning: Large Language Models (LLMs)
  • Class_Notes: Parallel Big Data Processing Systems on Cloud: NOSQL/NewSQL Systems
  • Class_Notes: Transaction and Concurrency Control




  • Class Announcement and POST



    10. Jan 14, 2026:

    The CIS612 Class Time Will Be 4:15PM - 5:30PM During This Spring Semester!




    Midterm Info:

    Information on Midterm and Midterm Lecture Links Only (Scroll Down to the Lecture Note Sections)


    The Topics to Focus On for the Midterm:

    IMPORTANT WARNING !!

    If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterm and the Final !!

    Characteristics and Architecture of Big Data Processing Systems/Applications, Cloud Computing Systems
    Universal Data Exchange Formats in Object Exchange Model(OEM): Three Common Big Data formats in Semi-Structured Model:
    Data Model, Syntax of Each Encoding Format, the Related Data Processing Techniques: DOM, XPath
    Characteristics of Semi-Structured Data Model, Main Differences between Semi-structed Data Model and Relational Data Model
    for Object Relation Mapping (ORM), Conversion among Relation(CSV), XML, and JSON

    Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining
    Comparison Between Semi-structured Database Server and Relational Database Server for Database Management Strategies


    NOTE that ONE Page Note is NOT allowed to the Midterm !!


    Exam Formats:

    6-7 Main Questions with 2-3 sub questions for a Small/Short Answer Types for Problem Solving





    Final Exam Information:

    Monday May 4th 4:00PM-6:00PM

    In Each Lecture Note for the Final, Only the Slides (Subjects) Covered in Class Will Be in the Final

    See the Lecture Notes Links Only for the Final here (All the rest of links are disabled)
    the Lecture Notes Links Only for the Final


    One Page Note (Each Side) Is Allowed to The Final. It Shoule Be Hand-Written. Printed or Copy and Paste of Entire Letcure Notes Are NOT Allowed.


    Topics to Focus on for the Final 

    Unstructured Data Processing, Inverted Index, How to Build Inverted Index, TF-IDF Based Ranking Algorithm with Cosine Similarity
    Text Cleaning Preprocessing, Basic Natural Language Processing (NLP) Methods in Data Pipeling, POS Tagging
    Problems of TF-IDF Based Ranking for Text Analysis
    Context Aware NLP Methods - POS Tagging, NER Tagging
    Design of Intelligent System with Data Pipelining of Big data processing

    Problems of TF-IDF Based Ranking for Text Analysis
    Context Aware NLP Methods - POS Tagging, NER Tagging, OpenIE
    Design of Intelligent System with Data Pipelining of Big data processing
    Information Extraction Methods, Design of an Answering System, How to Build a KnowledgeBase from a Large Collection of Unstructured Texts for Question Answering System
    Semi-Structured Query Model - XQuery
    Semi-Supervised Learning - word2vec for Question-Answering System
    XQuery Using XPath
    Intro to RDF Concepts
    Graph Database Neo4j
    Cypher Query Basics

    Parallel Map Reduce Programming, Parallel Distributed Data Processing Execution Algorithm in MapReduce (the Execution Phases (Steps) of Map Reduce on HDFS), Data Excution Algorithm Steps in Map phase and Reduce phase,
    Architecture of HDFS -- Architecture of Parallel Distributed File System in General
    Architecture of Parallel Distributed  Platform (Amazon EC2 Cloud Platform in General)
    Hive DDL to Build DW like Hierarchical Directory Structures
    Pig Latin Data Pipelining Execution Steps for General Analytic Queries 



    One Page Hand Written Note (Each Side) is Allowed to the Final !
    Printed Copy and Pasted Lecture Notes Are NOT Allowed as One Page Note









    9. Jan 12, 2026:

    Check the Final exam schedule and the Semester Schedule below
    University's Official Academic Calendar for the Semester Schedules to Add, Drop, Withdraw, and the Final Exam schedules


    8. Jan 12, 2026:

    It is Required to Attend Every Class !
    Your Class Attendance Will be Checked in Each Class.
    There Will be Random Quizzes to Check the Attendance !
    Only 5 - 10 Mins Will Be Allowed for Each Quiz. Those Who Come in Class Late After 5 Mins, They Will NOT Be Given a Quiz !  


    7. Jan 12, 2026:

    The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class Blackboard !
    Wait For the Lab Submission Links To Be Created on Blackboard with the Deadline to Submit Each Lab


    6. Jan 12, 2026:

    Only the registered students can access the course blackboard.
    If you have a problem with your blackboard access, please contact the registrar and CSU e-learning Tech Support to resolve the issue !

    Faculty do not control your registration in the CSU Campus Systems and the course blackboard access.

    e-learning CSU Tech Support




    5. Jan 12, 2026:

    TA Information:

    TA:
  • Kim Loc Chau

  • Email: k.chau@vikes.csuohio.edu

    Office Hours: Mon, Wed 1:00-3:00 PM (Officially); Tue, Thus 2:00-4:00 PM (if needed)

    Send email to TA ahead to set up a time slot to let him know that you are coming

    Location: Big Data Lab: FH 305 or a ZOOM Meeting

    ZOOM Meeting ID: 758 582 3674
    Passcode:



    If you have questions in the Labs or Grading your Labs, Send an email to TA to See During the TA's Office Hours or Schedule a meeting in the TA's Zoom meeting





    Dr. Chung's Office Hours:

    Tues and Thurs 1:30PM - 3:30PM
    Send Me Email Ahead to Set Up an In-Person or Zoom Meeting.

    Email: s.chung@csuohio.edu

    Zoom Meeting Info:
    Meeting ID: 859 3867 6332







    4. Jan 12, 2026:
    Lab Submission:

    The Output of each lab is your report in Doc file that shows your screen captures of the Executions with Each Output (each intermediate output as well as final outputs) Generated.
    your Report in Doc file also should explain all the platform set up, the execution steps, and copy of each source code files ) Each of your screen capture must show your results returned in YOUR SYSTEM to prove that your lab is done correctly by YOU !!

    Submit your Zip file that includes the followings on Blackboard for a timestamp and as a proof
    1) Your Lab Report in .doc file that explains all the platform set up procedures, the execution steps, each intermediate output, final outputs, and a copy of each source code files,
    2) All your Source files, and output files


    3. Jan 12, 2026:
    Each Class Attendance is Required for CIS612 !
    There Will Be Random Quizzes and Sign Up Sheet to Check Each Attendance of You !

    2. Jan 12, 2026:
    You can use any SQL Server for Your Database such as MySql or MS SQL Server. See Set up Guide for MS SQL Server or MySQL in the Lab Section.


    1. Jan 12, 2026:
    The class webpage for this semester was announced on the class blackboard:
    CIS612 Big Data Class Webpage


    Check the Last Day to Add and Drop here !
    University's Official Academic Calendar for the Semester and the Final Exam schedules

    CIS612Final Exam:
    Monday May 4 4:00p-6:00p







    Midterm Info:

    Information on Midterm and Midterm Lecture Links Only (Scroll Down to the Lecture Note Sections)


    The Topics to Focus On for the Midterm:

    IMPORTANT WARNING !!

    If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterm and the Final !!

    Characteristics and Architecture of Big Data Processing Systems/Applications, Cloud Computing Systems
    Universal Data Exchange Formats in Object Exchange Model(OEM): Three Common Big Data formats in Semi-Structured Model:
    Data Model, Syntax of Each Encoding Format, the Related Data Processing Techniques: DOM, XPath
    Characteristics of Semi-Structured Data Model, Main Differences between Semi-structed Data Model and Relational Data Model
    for Object Relation Mapping (ORM), Conversion among Relation(CSV), XML, and JSON

    Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining
    Comparison Between Semi-structured Database Server and Relational Database Server for Database Management Strategies.


    NOTE that ONE Page Note is NOT allowed to the Midterm !!


    Exam Formats:

    6-7 Main Questions with 2-3 sub questions for a Small/Short Answer Types for Problem Solving







    Final Exam Info for 2026:

    Monday May 4 4:00p-6:00p

    In Each Lecture Note for the Final, Only the Slides (Subjects) Covered in Class Will Be in the Final

    See the Lecture Notes Links Only for the Final here (All the rest of links are disabled)
    the Lecture Notes Links Only for the Final


    One Page Note (Each Side) Is Allowed to The Final. It Shoule Be Hand-Written. Printed or Copy and Paste of Entire Letcure Notes Are NOT Allowed.


    Topics to Focus on for the Final (Tentative)

    Unstructured Data Processing, Inverted Index, How to Build Inverted Index, TF-IDF Based Ranking Algorithm with Cosine Similarity
    Text Cleaning Preprocessing, Basic Natural Language Processing (NLP) Methods in Data Pipeline, POS Tagging
    Problems of TF-IDF Based Ranking for Text Analysis
    Context Aware NLP Methods - POS Tagging, NER Tagging
    Design of Intelligent System with Data Pipelining of Big data processing

    Semi-Structured Query Model - XQuery
    Information Extraction Methods, How to Build a Knowledge Base from a Large Collection of Unstructured Texts for Question Answering System
    Semi-Supervised Learning - word2vec and Question-Answering System
    XQuery Using XPath
    Intro to RDF Concepts
    Graph Database Neo4j
    Cypher Query Basics

    Parallel Map Reduce Programming, Parallel Distributed Data Processing Execution Algorithm in MapReduce (the Execution Phases (Steps) of Map Reduce on HDFS), Data Execution Algorithm Steps in Map phase and Reduce phase,
    Architecture of HDFS -- Architecture of Parallel Distributed File System in General
    Architecture of Parallel Distributed  Platform (Amazon EC2 Cloud Platform in General)
    MongoDB Queries: Aggregate Pipelining, Look Up Join Operator,
    Hive DDL to Build DW like Data Cubes, Mapping from HiveQL to Map Reduce Jobs,
    Pig Latin Data Pipelining Execution Steps for General Analytic Queries 
    Join Algorithms on a Parallel Distributed System in Map Reduce 

    Data Models of NoSQL Systems: Hive, PigLatin, MongoDB, Google Big Table/HBase (Not for This Semester)
    Transaction (Random Reads/Writes) for Concurrency Control of NoSQL Systems: Hive, MongoDB, Google Big Table/HBase
    Spark Data Processing Architecture
      Kafka Data Processing Models/Architecture (Not For This Semester)

    One Page Hand Written Note (Each Side) is Allowed to the Final !
    Printed Copy and Pasted Lecture Notes Are NOT Allowed as One Page Note







  • CIS612 Syllabus



  • Project and Presentation



    Projects on Big Data Processing, Training Large Language Models with Big data, Building an AI Application with Big Data Analytics


    Project Guideline and Important dates !
    Project Task Timeline

    Important Dates for Project: (Tentative)
    Proposal Submission by April 12 !
    Status Report Submission by April 21 !
    Presentation Starts on April 27 (Mon) and 29 (Wed) !
    Final Project Report by Saturday May 1 !


  • Read the Instructions of Project Presentation Here ! (The One drive Project Presentation Sign Up Sheet for Scheduler Will Be Sent To Your CSU Email !)

    Important Notes for Final Project:

    1. You Can Do One Person Group Project.

    2. 3-4 Person Group is Allowed with a Bigger Scale Project

    3. If You Need to Find Group Members, You Can Use the Blackboard Class Email to Send to the Class, Some will respond to you if they are looking for a group member.

    4. Implementing Your Final Project with Clientside and Serverside of a Web Application with a Web Based User Interface Is NOT Required.
    (It will Be Counted More If AI/Big Data Project is impleneted as a Web Application (Created as a Website with a Webserver).

    5. You can Change Your Project Plan/Proposal or the Details of Proposed Project even after Your Proposal has been submitted until the Deadline of the Status Report.

    6. For Those Who Have Aleady Taken CIS660, it is Required to Complete a Full Project with Bid Data Processing and Training ML to Have Predictive Modeling


    1 - 2 Person Group Project Are Allowed for the Small Class Size (~25 students);
    1 - 3 Person Group Project Are Allowed for the Big Class Size (> 35 students);
    1 - 4 Person Group Project Are Allowed for the Very Big Class Size (> 40 students)
    If 3-4 Person Group is Allowed, Make Sure to Make the Project in a Bigger Scale

    If 3-4 Person Project Group Will Be Allowed As Long As the Project Scale is Big Enough for 3-4 Persons

    PhD Student Should Do One-Person Project (Not in a Group)





    Final Group Project Specification and Instructions:

    Project Submission Instructions:

    One Submission Per Group Required.

    Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week !
    Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx !

    If your data file is too big to upload, Submit your zip file with your Data file either on the google drive or One Drive and send the link. (Send me email for access permission for this link !)



    Submit a Zip file that includes:

    - All of your Group Presentation Slides (Must be .ppt)
    - Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files on Blackboard by the end of Friday of your presentation week.
    - Include the Problems/Error Encountered and Your Resolutions in Your Report
    - Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files.
    - The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results.
    - If Your Group Report/Presentation Don't Show/Include Any of the Required Contents in Your Report and Presentation,
    I will ASSUME that your group submitted a Copy of Somebody's Github Codes that were downloaded from the Internet.




  • Project Specification and Instructions on What To Submit
  • Project Task Description and Suggested Project List (Will be Added More)
  • FAQs for Final Project








  • Grading Criteria of Project Complexity Based on:

    - Big Data Size,
    - Complexity of Big Data Processing/Transformation Methods
    - Complex Big Data Processing Methods or Analytic Scoring/Ranking Methods
    - Superior Knowledgebase Design Such as Grapgh or Property Based Complex Knowdlege Representation Design

    (Naive Structured CSV/TSV Data or Already Processed File Downloaded from Kaggle sites Will Not Be Considered as Big data)

    Preferred Big Data Project Platforms:
    They are not required but Will Be considered more !

    1. The Platform/Systems with Any Parallel Database Server with MapReduce and Hadoop Distributed File Sysrem
    Or 2. Real-time Big Data Analytic Processing with Spark 




    Some Suggested Real life Big Data Sets for Group Projects

    JSON Data sets at Yelp Site

    Social Media Twitter Server Generated Log Stream Data in JSON:

    Social Media Twitter Data in JSON: 2020 Presidential Election Candidate Trump and Biden Data Sets

    Social Media Twitter Data in JSON: farmers protest twitter data set


    Medline Plus for Medical Encyclopedia

    7.2 Million Wiki Page Dump either in HTML or XML

    7.2 million wiki webpage dump in XML

    32 million PubMed Collection of Biomedical Research Paper Abstracts in XML

    Any document collection of clinical notes or doctors’ notes
    MIMIC: Medical Patient Data to get Doctors’ Text Notes

    The IMDB Data links in the Project List are gone. Check here for Movie Review Data Set.
    Movie Review Data

    Youtube Analytic Site




    Examples of Good Final Big Data Projects

  • Example of Final Project: Building Medical QA System on Medline Plus Site
  • Example of Final Project: Building a Prediction System from Social Media Data for 2020 US Presidental Election




  • For Those Who Have Already Taken CIS660, The Final Group Project Should Include Fully Analytic Processing

    Suggested Projects:

    For Text Analytics like Sentiment Analysis or Opinion Analysis: NLP Techniques - POS, NER Tagging, Bi-Gram Handling Are Required for Preprocessing.
    For Document Categorization: by Constructing TF-IDF Vectorization. Inverted Index Building Will Be Plus but Optional.
    Building Word2Vec Embeddings for a Collection of Documents/Webpages with Training Set Generation in Skip Gram Model
    For Other Types of Projects, the Proposal is Required to Be Approved to Meet the Complexity of Final Project


  • Requirement for Final Project Report (See The One Drive Presentation Schedule Sheet Sent Out or the Final Submission Instruction in the Project Section
  • Group Project Presentation Schedule (Invitation to One Drive Sign Up Sheet for Presentation Schedule Sheet Will Be Sent to your CSU CampusNet Email)

  • See the Project Page for Common Project Ideas or Examples of Big Data Projects

    Group Project ***********************

    Project List is posted.
    Project List






    IMPORTANT Submissions for Group Project

    Task 1: Group Project Proposal (Plan)
    Submit minimum 3 Page Group Proposal on Blackboard with List of Group Memebers

    Group Project Proposal (Plan) Should Include Brief Descriptions on:
    1. Description of Big Data in Size and Format, Data Collection Plan, 2. Goal of Your Big Data Analytic Project with What Kind of Intelligent Analytic Funtionality, Features of Your AI/Big Data Analytic Application
    3. Big Data Processing Plan, Methods
    4. Investigate on Platform/Systems/Tools/APIs to Use


    Task 2: Group Project Status Report

    Your Project Status Report Should Show the Following Tasks Done:
    1. Your Big Data Collection Done 
    2. Platform Setting/System Configuration Procedure (if it is new)
    3. Design of Big Data Processing Pipeline, Data Transformation Methods
    4.Your KnowledgeBase Structure/ Database Design
    5. Your Applucation/Analytic Server Codes, Processed Data/Database Contents in Progress, Any Intermediate Outputs in Progress



    Task 3: Project Presentation Project Presentation Starts from the Last Week of the Class of the Semester
    Project Presentation Schedule Will Be Sent To Your CSU Email for Sign Up One week Before the Presentation
    Read the Instructions of Project Presentation Here !


    Group Project Presentation Should Include:

    1. Data Description, Data Size, Data Collection Method
    2. Goal of Your Intelligent Big Data Analytic Application (AI)
    3. Platform Setting/System Configuration Procedures
    4. System Design (Architecture) of Your AI Application in Detail
    5. Raw Big Data Preprocessing Methods and Intermediate Results
    6. Design of Big Data Processing Pipeline, Data Transformation Methods
    7. Description of Your KnowledgeBase Structure/Database Design and Real-time Question to Query Processing for Your AI Application. Show the Contents
    8. Your ML Algorithms if any, Ranking Algorithm if any for Labelling if any
    9. Your ML Models if Any and Evaluation (Accuracy) Results and Visualization of the Result, Data Marix (Structure) for Analysis/Evaluation and Visulation of the Analysis Results (for example, correlation matrix or similarity matrix) if Any
    10. The Problems/Errors Encountered and Your Resolutions
    11. System Demo




    Submit Final Project Report and Presentation By the End of Friday of The Last Class Week After Your Presentation!

    One Submission Per Group Required.
    Submit a Zip file that includes:
    1. All of your presentation slides (both in .pptx) and
    2. Your Group Final Project Report (in doc) 
    
    The Final Project Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results.
    
    The Report Should Include Platform/System Set up the Set Up Procedure /Configuration Detail of Your Platform/System/Packages, Executions Steps, all the Source Codes, Scripts, all the intermediate outputs, and final output files.
    Include the Problems/Error Encountered and Your Resolutions in Your Report
    
    If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web.
    
    Group Project Presentation Should Include:
    
    1. Data Description, Data Size, Data Collection Method
    2. Platform Setting/System Configuration Procedures
    3. System Design (Architecture) of Your AI Application or Data Analytic Goal in Detail
    4. Raw Big Data Preprocessing Methods and Intermediate Results
    5. Design of Big Data Processing Pipeline, Data Transformation Methods 6. Description of Your KnowledgeBase Structure/Database Design. And Show the Contents
    7. Ranking Algorithm, Data Matrix (Structures) if any for Evaluation
    8. The Problems/Errors Encountered and Your Resolutions
    9. System Demo or Evaluation Results and Visualization of the Result
    















    Project Examples of Vectorization of each document in a Training set for a Machine Learning Classifier:

    Project Example I: Sentiment Analysis of Yelp Business Review with Machine Learning

    Project Example II: Sentiment Analysis of Yelp Business Review with Machine Learning

  • Wikipedia Search Engine by Nick McCoy, et al
  • Word2Vec and Paragraph Vector for Document Clustering by Mike D'Arcy and Utkarsh Patel


  • Best Group Projects: Selected Best Projects Will be Posted Here !

    Big Data and Data Science Projects:

  • Social Media Opinion Analysis System for 2020 Presidential Election Prediction (the Candidate of Best Senior Project in Engineering College of 2021 (Created From the Final Project of CIS408 and CIS430)
  • 2021 Senior Project: Stock Market Analysis Service System
  • Intelligent Infectious Disease Tracking System
  • Product Review Sentiment Analysis System

  • Sharesci: Online Research Document Search Engine in Natural Language (with Mongo DB, Node Js and Angular JS) by Mike D'Arcy and Utkarsh Patel





  • More To Come Here !


    Group Project Data Sources: You can choose to work on these data sets for your group project

  • Wiki Page Data set
  • Complete Yelp Callenge Data set: 5 BIG JSON Files
  • More to Come Here !


    Final Project Submission Instructions:

    Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week !
    Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx !
    If your data file is too big to upload, Submit your zip file with your Data file on your google drive or One Drive and Send email to me and TA to share !

    Submit a Zip file on Blackboard by the end of Friday of your presentation week.
    One Submission Per Group Required.

    Your Project Zip File Should includes:

    1) All of your presentation slides (both in .ppt and .pdf) and
    2) Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files
    3)Include the Problems/Error Encountered and Your Resolutions in Your Report

    Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files.
    The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results.

    IMPORTANT NOTE !!!
    If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web.


  • Lab Assignments and Resources



    Basic Python Tutorials :

  • Python Tutorial
  • Python Codecademy Tutorial Site



  • Guides for Important Platform Set up with Python, XPath, Beautiful Soup, and MySQL:

    Platforms Set up Guides for Big Data Labs (by TA Dennis Risch) *******************************

    See More Guides in the Lab1 Section (Far Below)



    Python Data Science Platforms:

    Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system)

    • Anaconda
    Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms
    Anaconda Tutorials

    • Python Anaconda Tutorial Sites
    Anaconda Tutorial Site
    Anaconda Tutorial Site
    • PyTorch
    PyTorch Site (It can be integrated from Anaconda as well)

    Google Colab:
  • Google Colab

  • Basic Guide for Python tools for Data Science: jupyter-notebook
    More Basic Guide for Python tools for Data Science: jupyter-notebook

    • Python Scikit Learn for Common Data Science Tasks
  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing
  • Python Text Processing Libs for Text Analysis
  • Python Numpy Tutorial

  • Text Preprocessing (Natural Language Processing) Library in Python SpaCy:

    Liquistic Modules in Python SpaCy
    Lemmatizer in Python SpaCy
    Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy
    How to Code Liquistic Modules like Lemmatizer in Python SpaCy
    Python Example for Basic Text Processing

    • Python sklearn.cluster Python Sklearn Clustering



    Python IDE Deduggers:

    Basic Guide for Python Debugger Pycharm
    Python IDE Spyder
    Python Debugger Spyder




    Basic R Tutorials :

  • R Studio basic Tutorial
  • R Basic Online Lecture

  • R Manuals

  • R Tutorial with Examples with R Stat Tool
  • Note that Examples in this tutorials may not the final correct output for Lab1 !
  • Examples of Basic R Stat Tool with Helpful References

  • Special Online Study Guides on Basics on Data Warehouse/OLAP, Data Analytics, Big Data in Independent Study Independent Study with Nick White (Now in FaceBook and The First Prize Winner of 2016 Senior Project)



    Useful Machine Learning Tutorial Sites:
    Keras for Image Processing/Text Processing with Deep Learning:
    Keras Machine Learning


    For Your Own Advanced Study
  • Coursera Machine Learning by Stanford
  • Open AI
  • Google Colab
  • Medium by MIT



  • Lab Submission Instructions:


  • The Output of each lab is your Lab Report in Doc file that shows your screen captures of each of your executions with Your Outputs

  • Your Report in Doc file should include all the platform set up procedures, the execution steps, and copy of each source code files
  • Each of your screen capture must show your results returned by your systems and your database servers to prove that you have done the lab correctly !!



  • 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof.

    2. IMPORTANT NOTE:
    Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System.

  • Example of Output of the Execution Steps for Labs

  • 3. If You did Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front Page of Your Report in Bigger and Bold Font !



    Useful Lab Helpers:

  • XHTML Validator
  • How to debug HTML, JAVASCRIPT, DOM, XPath
  • JavaScript Debugger
  • Node.JS Setup Recommendation
  • MDN Site for DOM with XPath in JavaScript
  • XPath Setup Guide for JavaScript
  • XPath Methods
  • Setting Up DOM with XPath for Python
  • Webscapping with DOM, XPath for Python with BeautifulSoup
  • Python XPath Guide

  • Useful Tools: Swagger API Tool for REST API Developments

    Useful Big Data Analytic Tools

    Choose your System/Tool/Platform to Set Up and Get Used to:

    Machine Learning in Python:

  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing
  • Python Numpy Tutorial

  • Natural Language Processing (NLP) for Text Preprocessing/Big Data Analytics Library in Python SpaCy:

  • Python Text Processing Libs for Text Analysis
  • Liquistic Modules in Python SpaCy
    Lemmatizer in Python SpaCy
    Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy
    How to Code Liquistic Modules like Lemmatizer in Python SpaCy
    Python Example for Basic Text Processing









    Lab Assignments:


    Instructions for Lab Submission:

    The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard !
    Wait For the Lab Submission Links Are Created on Blackboard for Each Lab


    You Have to Start Working on Labs Before the Submission Link Are Created on Blackboard for Each Lab Submission.

    Always Follow the Deadline of Each Lab Assigned on the Class Blackboard.
    The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester.

    Please Identify Your Course When You Ask Me in Email !






    Each of your screen capture must show your results returned by your systems and your database servers to prove that you have done the lab correctly !!



    1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof.

    2. IMPORTANT NOTE:
    Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System.

  • Example of Output of the Execution Steps for Labs

  • 3. If You did Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front Page of Your Report in Bigger and Bold Font !










    Lab0: Learning Python -- Due by the End of the Second Friday of the Semester

  • Python Codecademy Tutorial Site

  • Python Coding Guidelines

  • Python Data Science Platforms: See More Python Platforms Above or Lab1 Set Up Guides Below to Choose for Data Science

    Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system)

  • Installation and Setting Up Either MySQL or MS SQL Server

  • Installation Guides for MySql Server:

    MySql Download
    MySql Download and Installation
    MySql Tutorial
    How to Create MySQL Database


    If You Want to Use MS SQL Server, Installation Guides for MS SQL Server:

    See the Announcement Section of CIS430/530 Database Systems and Processing below for Account Creation for Microsoft Azure site for Free Download of MS Visual Studio and MS SQL Sever.
    How to Create MS Azure Potal Site Account

    See the Lab Section of CIS430/530 for Installation Instruction of MS Visual Studio and MS SQL Sever.
    How to Download and Install MS SQL Server


















    XML/XHTML/JSON Syntax Validators:

  • XHTML Validator
  • JSON Validator
  • How to debug HTML, JavaScript, DOM, XPath in a Web Browser













  • Lab 1 on Web Data Processing for Information Extraction

  • Lab1 on Webpage Processing for Information Extraction
  • FAQs for Lab1 on Information Extraction from Webpages


  • Infoplease site of State Union Addresses of US Presidents
    Correct page of Address of John Adams December 3 1799


    The Simplied html file of the Info site (You may use this simplified html file to inspect the element structure to extract the required info for Lab1)


    Important Notes:

    Infoplease site of State Union Addresses of US Presidents (As of 2025, This site blocks any http request from unknown clients (you))


    You can either using the workaround below to avoid blocking or use another site for State Union Addresses of US Presidents

  • How to Avoid Sever Blocking http file Request (Shared by Collin Simpson and TA Kim Loc Chau)



  • Another site for State Union Addresses of US Presidents that Does NOT block http request

  • Another site of State Union Addresses of US Presidents


  • Implementation Related Important Notes:

    If there is a Link that Does NOT Have Any Web Page Contents, Add NULL Values for the corresponding Columns for the link

    Do not Assume that every sites has an identical URL format. This is semi-structured data. Nothing is regular in Big data.
    For the irregular parts, use regular expressions or xpath as neccessary.




    For Those Who Took CIS593 Big Data:

    Lab3 on Information Extraction from Customer Reviews on the Social Media Websites Using DOM and XPath

    Visit Yelp Cleveland site below to Extract each customer review in each restaurant page.

    Yelp Site for Best Restaurants in Cleveland, OH

    For example, Extract each customer review in the restaurant page to store them in a SQL database table
    Website for a Restaurant with Customer Reviews

    Create a table with Restaurant name, Location(Address), Reviewer Name, The number of stars, Review Text


    Make it Automated for the Table Creation from your output file (CSV file) in your Database Server !






    General Approaches for Webpage Processing

    There are two ways to do Information Extraction from Webpages.
    Method 1 with HTML DOM and XPATH is REQUIRED for Lab1 !!!

    Method 1: Webpage as Semi-Structured HTML DOM Tree Using XPATH (More Scores Given!)
    Method 2: Webpage as Unstructured Text Using Parsing API like Beautiful Soup





    Set Up Guides for Important Platform with Python, XPath, Beautiful Soup, and MySQL:******************************

    You Need To Set UP DOM and XPATH for Any Server-side Scripts/Languages for Applications such as Lab1.


    Platforms Set up Guides with Python, MySql, XPath, PyCharm Debugging IDE for Big Data Labs (by TA Dennis Risch) *******************************


    More on Setup Guides: Python Setup Guides for Labs:

    pip Python package Installation Guide
    Setup Guide for Anaconda Python Framework and Jupyter Notebook IDE (by TA Hemal Paneliya)
    Setup Guide for Anaconda Python Framework and Spider IDE and Debugger (by TA Durga Dasepalli)

    pyodbc Set Up Guide With MS SQL Server:
    How to Set Up Python Jupyter Notebook to Connect to a Microsoft SQL Server using pyodbc****************


    Platform Set Up Guide on Mac:
    Installation Guide on Mac for lxml with pip Python package for DOM with XPath and pycharm IDE for debugging on Mac

    /

    More on Python IDE Deduggers:

    Basic Guide for Python Debugger Pycharm
    Python IDE Spyder
    Python Debugger Spyder


    Database Server Set up is needed for the Labs.
    You can use any SQL Server -- MySQL, MS SQL Server, or any Database Server

    See Lab0 Section Above for more instructions or See the step-by-step installation guides in CIS430/530 Lab Section Below
    CIS430/530 Lab Section
    How to Download and Install MS SQL Server



    Note that the Set Up Guides Above are for Server Side Applications such as Lab1 in Python or Java.
    You don't Need to Do Any Extra Set Up to Use XPATH in a JavaScript in a Client side Codes for HTML DOM Processing which will be executed by your Web browser.
    All the Recent Web Browsers Have XPATH Features in their Debugger by Default.








    IMPORATNT NOTE !!!!

    The tutorials and examples in the class Lab Section are to guide and teach the methods and techniques.

    Please note that they are NOT to provide you with precise coding solutions that would run without checking/debugging the changes in codes;
    Especially since these areas are fast changing and there are so many variations in the platforms/languages.
    YOU ARE RESPONSIBLE for DEBUGGING YOUR CODES !!



    Lab1 Implementation Guides:*******************************

    Lab1 Guides: What to Need to Know in the Lab1 Section to Do Lab1 *************************

    Note that the Code Examples of the Lab1 Guides below Do NOT Contain the Complete Codes to Be Executed. These are ONLY for the guides for Lab1.
    The Environment Configuration and the Versions of APIs Varies.
    Do Not Copy the Entire Codes Blindly to Do Your Lab1 since it will not work depending on the version of your python and setting up for DOM and XPath.

    How to Write Xpath in Demo in Webbrowser:
    Example of Xpath Execution in Webbrowser *************************


    Code Examples with DOM and XPATH:
    Example of Lab1: General Guide in Python with DOM and XPath, and pyodbc for database operations for Information Extraction *************************
    Example of Lab1: Guide in Python with DOM and XPath in lxml for Information Extraction *************************
    Example of Lab1: Guide in PHP with DOM and XPath for Webscapping
    Example of Lab1: Guide in Python on Linux for Webscapping

    Code Examples with Beautiful Soup Text Processing API:
    Example of Lab1: Guide in Python with Beautiful Soup for Information Extraction



    DOM with XPATH Set Up in Any Script/Programming Languages for Server Side Applications
    Note that You don't Need to Do Any Extra Set Up to Use XPATH in your Serverside JavaScript in NodeJS.


    All the Labs of CIS612 Are Server-side data processing in a server-side script/programming language with file I/O and DOM parser with Xpath.
    They are NOT a client side JavaScript executed by your Web browser.

    For the Labs to Write an Application Server in this course, You Need To Add DOM Parser related APIs for Your Scripts/Programming Languages To Do the following steps:
    to Call a DOM Parser API
    to Build a DOM Tree and
    to Use XPATH Methods for Retrieval to Extract information You need as in the Examples below.

    Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use.

    Setting Up lxml parser for DOM with XPath for Python *****************************
    Node.JS XPath Setup Guide and XML Parser in JavaScript based Node JS
    XPath Setup Guide for JavaScript
    XPath Setup with npm in JavaScript for Node JS


    Set Up DOM with XPath for HTML and XML:

    How to Debug HTML with JavaScript, DOM, XPath
    Learn How to Inspect Element in Chrome to Debug HTML with for DOM and XPath

    JavaScript:
    JavaScript Debugger
    XPath Examples in Javascript
    MDN Site for DOM with XPath in JavaScript
    DOM with XPath Setup Guide for JavaScript
    XPath Methods

    Node.JS XPath Setup Guide
    npm Node.JS XPath Setup Guide
    Google Crome XPath Helper Setup

    Python:
    Setting Up DOM with XPath for Python
    BeautifulSoup Documentation for HTML, XML for Python
    Webscapping with DOM, XPath for Python with BeautifulSoup (From the Stanford Class)
    Python XPath Guide


    Automatic Table Creation in a SQL Server

    After installing a SQL server (See the step-by-step installation guides in CIS430/530 Lab Section
    Add the codes in pyODBC as below for a Connection for the Anaconda Python Framework to Connect to Your SQL Database Server

    For a connection with SQL Server with servername and database name using pyodbc in Python:

    conn = pyodbc.connect('Driver={SQL Server};Server=YOUR_SQL_SERVERNAME\SQLEXPRESS;Database=YOUR_DATABASE_NAME;Trusted_Connection=yes;')


    pyodbc(Open Database Connectivity) to Connect a MySql Server in Python:

    How to Connect to MySql Server with pyodbc in Python
    Stackoverflow site for pyodbc: Examples to Connect to MySQL


    pyodbc(Open Database Connectivity) to Connect MS SQL Server in Python:

    How to Connect to MS SQL Server in Python
    How to Connect to MS SQL Server in Python
    pyodbc: Examples of ODBC in Python
    Installation pyodbc to Connect to MS SQL Server in Python
    Stackoverflow site for pyodbc to Connect to MS SQL Server in Python


    Column Data Types to store a Large Text data to Create a Table with in MS SQL Server or other database server

    How to Create a Table in a SQL Server from CSV/TSV Text files

    How to Create a Table from a File with Bulk Insert with MS SQL Server
    How to Create a Table from a file with Bulk Insert in MySQL Server
    Example of a Script to Create Multiple Tables from different files with Bulk Insert in MS SQL Server



    Earlier Features to Handle Big Data in Relational Database Server

    Advanced Data Types: BLOB (Binary Large Object) or Text/CLOB (Character Large Object) in MySQL or MS SQL Server

    How to Create and Insert to a Table with a Column of Large Text Data or Image Data in a Relational Database Server:

    Text Data Type in MySQL
    What is TEXT data type in MySQL
    Difference between blob and clob datatypes
    Blob Data type in MySQL
    how to insert blob and clob from files in mysql
    how to insert text column in mysql



  • Trouble Shooting Error Resolutions When Set Up SQL Server with ODBC/JDBC

  • For Those Who Want to Use CLR Table Function to Create a Table in SQL Server -- This is NOT For CIS492/593
    How to Set Up ASP.NET with SQL Server
    How to Debug CLR UDF, CLR UDT, CLR TVF




    How to Connect to MS SQL Server in Python
    pyodbc: Examples of ODBC in Python


    Automatic Table Creation in a SQL Server

    pyodbc(Open Database Connectivity) to Connect MS SQL Server in Python:

    Installation pyodbc to Connect to MS SQL Server in Python
    pyodbc to Connect to MS SQL Server in Python
    How to Connect to MS SQL Server in Python
    pyodbc: Examples of ODBC in Python

    How to Create a Table in a SQL Server from CSV/TSV file

    How to Create a Table with Bulk Insert with MS SQL Server
    Example of a Script to Create Multiple Tables with Bulk Insert in MS SQL Server



  • Trouble Shooting Error Resolutions When Set Up SQL Server with ODBC/JDBC















  • Lab3_2:

  • XML Data Processing with DOM and XPath for (ORM) Object Relation Mapping:

    Lab3_2 on XML Processing with DOM and XPath for ORM
    Lab Assignment 3_2 - XML Input File
  • Transformation Rules from Semi-Structured Data Model (XML or JSON) to a Relational Scheme in the Third Normal Form (From the CIS530 Database System Lecture)

  • General FAQs on Semi-Structured Data Transformation to a Relational Scheme

    Check the validity of XML Input. Add a proper XML doctype heading like in your XML input document and Validate your XML input and correct it if needed.

    LabAssignment 3 - DOM XML Parser Example Codes
    Lab3 Output Examples of Conversion Between XML and SQL Table
    Example of Stored procedure to Create SQL Tables from a Text file


    JAVA based XML DOM Parser to download to set up if you are using JAVA for this lab.

    DOM XML Parser
    SAX XML Parser

    Examples Conversion Between XML and Table

    FAQ on Lab3_XML and Lab3_JSON









    Lab 3_3 on ORM Mapping with Semi-structured JSON Data Processing

  • JSON Data Processing:


  • Lab3_3: JSON Data Transformation to Relational Table
  • FAQs on Lab3_3 on JSON Transformation
  • Transformation of JSON to a Correct Relational Scheme in the Third Normal Form as well as the First Normal Form (From the CIS530 Database System Lecture)

  • Notes and Corrections:

    Note that you need to transform business.json file only for Lab3_3, not all of 5 JSON files from the Yelp site.

    Yelp Business Data Set for Lab3_3:

    Data Sets for Lab3_3:

    The given OneBusiness.json file and business100.json file in JSONData.zip are the corrected files.
  • JSONDATA zip file


  • Yelp Full Data Set Also Available in the Big data Lab below:

    Zip file for 5 JSON files from Yelp data Challenge 2017 or Download directly from the Yelp site below for 2020 data sets!

    Note:
    If You are having a corrupted zip file error for 2017 Yelp Data Set, use 7-Zip to unzip.


    You can directly download the most recent data sets from Yelp site below



    Yelp Data Set and Documentation for JSON File Structures

  • Yelp Site for Full Data Sets for Big Data Project
  • Yelp site documentation for JSON File Structures
  • Yelp site for Data Description


  • Note:
    Those Full Data Sets May Not Be in a Valid JSON Format. (Always expect this in a real world project)

    You Have to Use 7-Zip to unzip the zip file. Some other compression software might not be able to unzip correctly.

    The full JSON data files from the Yelp site might have a few incorrect JSON syntax detected in the data file or invalid line feed, which is common. Correct them before processing if detected.

    See FAQs for how to correct invalid JSON data

    Invalid JSON format handling:
    If there is any invalid data format is found in the input file, you can change it to the correct JSON format. For example, $$ in the “Price Range” key value pair in your input json file, the value $$ is not in quotes and this will cause to fail.

    Suggested Solutions:
    Replace $$ with 2 (meaning the price level is 2 in scale 1 - 5) in the file and try to parse the corrected file in your program.
    If , is missing between objects, add it.


    If you have invalid JSON file problem, the JSON file might need to be corrected through the JSON validator to detect the errors first then correct the JSON syntax to be imported. The data files in the Zip file are all coming from the yelp site. They sometime have the incorrect JSON syntax, which need to be corrected to make it work. 
    See the example of how to correct an invalid json file below. Scroll down for the part. FAQs for JSON Processing in General

    More FAQs on Lab3_3
    Q:
    Can we use the 2019 dataset (JSON) from yelp website for our Lab 3_3? Or we should only use 2017 JSON data?

    A:
    Either data set is ok. Some people have a Zip error with the 2017 data set. If then use the 2019 data set from the Yelp site.











    Lab 3_4 on Semi-structured JSON Data Processing with MongoDB and ORM

  • JSON Data Processing with MongoDB and ORM (Object Relation Mapping)

    Creating Semi-Structured Database: MongoDB Collection in a MongoDB Server from business.json and review.json files from the Yelp site No

    Lab3_4 on Creating Collections from JSON data files of Yelp Business and Review Data (NOT For This Semester !)



    Data Sets for Lab3_3 and Lab3_4:

    The given OneBusiness.json file and business100.json file in JSONData.zip are the corrected files.
  • JSONDATA zip file


  • All Yelp Data Sets (zip file)


  • Yelp Full Data Set Also Available below:

    Zip file for 5 JSON files from Yelp data Challenge 2017 or Download directly from the Yelp site below for 2020 data sets!

    Notes:
    - There are differences between 2017 Business Data and 2020 Business Data. Either Data Set is OK for Lab3_3 and Lab3_4.
    - If You are having a corrupted zip file error for 2017 Yelp Data Set, use 7-Zip to unzip.


    You can directly download the most recent data sets from Yelp site below



    Yelp Data Set and Documentation for JSON File Structures

  • Yelp Site for Full Data Sets for Big Data Project



















  • Lab 3_4:

    NoSQL Database System MongoDB with Data Aggregate Pipelining for Information Extraction and Data Analysis

    Lab3_4 on NoSQL Database System MongoDB with Data Pipelining for JSON data processing


    Example Runs of Mongo DB Queries


    Create Semi-Structured Database in MongoDB for two JSON data files: Business.json and Review.json from the Yelp data site

    1. Import Business.json and Review.json Data files from Yelp site into MongoDB
    2. Retrieve Information Using MongoDB Aggregation Pipelining for Information Extraction



    Note that:


    You have to use a full business.json data and a full review.json data file.
    If you use 100 business data only, you will get an empty query result.
    If Q2_1 or Q2_2 returns an EMPTY result, Change the Filtering Condition review_count Until You Get the Reasonably Good Number of Results.
    For example,
    To avoid the empty query results, the filtering conditions have been changed to review_count > 5 and star <= 2.
    For Q2_1: the review_count > 5
    For Q2_2: the review_count > 5



    Note for MongoDB Join for Part2:

    If You are using 100 business data for the join opeartion, there are no matching business id in the business ids from the list of 100 businesses and the review collection, so join results were empty.

    The reasons for this could be either 100 business data is too small to get any join result with the review data or the review data set in the new data set from Yelp 2019 has been changed. They keep changing or updating the Review data sets,
    so there is no matching id in the 100 business data set in the class web page which was collected in 2017.

    Use the entire Business and Review json files in the new data sets from 2019 to have meaningful join results.



    Note that You have to process the entire data set in each json file, not just 100 Business data.
    Make Sure to Use All the Documents in business.json to create a Collection named business and another Collection named review for this Lab


    Yelp Full Data Set Available in the Big data Lab below:

    Zip file for 5 JSON files from Yelp data Challenge 2017 or Download directly from the Yelp site below for 2020 data sets!

    You can directly download the most recent data sets from Yelp site below

    Yelp Data Set and Documentation for JSON File Structures

  • Yelp Site for Full Data Sets for Big Data Project
  • Yelp site documentation for JSON File Structures
  • Yelp site for Data Description


    The given OneBusiness.json file and business100.json file in JSONData.zip given in Lab3_3 are the corrected files.
  • JSONDATA zip file




  • Important Notes:

    If You are having a corrupted zip file error for 2017 Yelp Data Set, use 7-Zip to unzip.
    You Have to Use 7-Zip to unzip the zip file. Some other compression software might not be able to unzip correctly.


    Those Full Data Sets May Not Be in a Valid JSON Format. (Always expect this in a real world project)

    The full JSON data files from the Yelp site might have a few incorrect JSON syntax detected in the data file or invalid line feed, which is common. Correct them before processing if detected.
    See FAQs in LAB3_3 SECTION ABOVE for how to correct invalid JSON data


    If you have invalid JSON file problem, the JSON file might need to be corrected through the JSON validator to detect the errors first then correct the JSON syntax to be imported. The data files in the Zip file are all coming from the yelp site. They sometime have the incorrect JSON syntax, which need to be corrected to make it work. 
    See the example of how to correct an invalid json file below. Scroll down for the part. FAQs for JSON Processing in General




    More FAQs on Lab3_4:
    Q:
    Can we use the 2019 dataset (JSON) from yelp website for our Lab 3_4? Or we should only use 2017 JSON data?
    And Do we need to create CSV for Q1 as well for the count result?

    A:
    Either data set is ok. Some people have a Zip error with the 2017 data set. If then use the 2019 data set from the Yelp site. From the results of Q2_1 and Q2_2 to extract the info (review_id, business_id, stars, review_text) to convert to CSV. Conversion is not required for the Q1 Result.

    Q: How many Extra Credit we will get if Lab3_4 is built as a full Rest API Application?
    A: It Will Be Counted 20% Extra Credit.




    How to Install pymongo driver and Connect to MongoDB Server in Python application as a client

    See MongoDB Lecture Notes Section for MongoDB CRUD Queries and more details

    How to Import Data file to MongoDB
    Mongo import


    Sample Runs of Mongo DB Queries

  • Sample Project Codes: Processing JSON Data in MongoDB to SQL Table in Python

  • How to Import a json file to MongoDB

    How to Use MongoDB in Python for CRUDE: DB/Collection Creation, Insert, and Querying

    How to Import a JSON file to MongoDB in Python

    Example of MongoDB Aggregation Pipeling for Word Count

    How to Save MongoDB Query Results into a variable

    How to Save MongoDB Query Results















    Lab3_4: EXTRA CREDIT (30%) !!!

    Collecting Real Time Twitter Message Streams by Topics and Creating Semi-Structured Database in MongoDB for Retrieval

    If you choose Twitter Stream Data Collection to Do Lab3_4 instead of Yelp Data, it will be 30% Extra Credit !

    Lab Specification:
    Lab3_4: Twitter Data Stream Collection, Collection in MongoDB, and JSON Transformation


    Twitter Logging Structure in JSON
    Note some Tweets Don't have the Retweet part of info. Try to Collect Retweets Together.


    Note that:
    Lab3_4 Report to do Real Time Twitter Stream Data Collection and Use of MongoDB to Store and Maintain Them.
    The Collected Twitter Data Will Be Used for Sentiment Analysis as Your Final Project later






    Most Recent Update on Twitter Free Account for Data Collection (By Collin Simpson)




    Twitter Data Sets:

    If You Can't Get Any Tweets Using Your Twitter Developer's Account, Use this Twitter Raw data set in this site:
    Raw JSON Twitter Data (zip) -- farmers-protest-tweets-2021-2-4

    Raw JSON Twitter Data (zip): coronatweets_11_53.zip

    Raw JSON Twitter Data (zip): Tweets on Trump and Biden (Collected One Month before the 2020 Presidential Election)

    Kaggle: Clean Raw JSON Tweets Data site



    Note That if TWITTER Real Stream Data Collection is not done, Extra Credit *will not be given !




    For Twitter Stream Data Collection:


    NOTE that to Collect the Twitter Real Time Stream Data, you Need to Apply for Their Developer’s Account in the Twitter Developer's site

    You Need to Get a Permission to Get credentials for a token and keys. This process usually takes 3-4 days or one week for Twitter to respond.
    You HAVE TO Start Your Twitter Application ASAP. Do NOT Wait until the Last day.
    See the example project and sample codes for the step by step procedures for this. Read everything posted in the links below.


    For Your Twitter Developer Account Application:
    Apply a Twitter Developer's Account ASAP for Twitter Stream Data Collection ! Start ASAP Since it Will Take a Week to Obtain a Twitter Developer's Account !!

    DO NOT Choose Elevated, Advanced, or Academic Research Developer's Account to Apply. You Might NOT Get a Permission and It Will take More Time

    FAQs to Get Twitter Developer's Account


    Twitter API 1.1 is depreciated and wont be available for new developers:
    Twitter API 1.1 is depreciated

    The New Structure for the 2.0 API for tweets: the New Structure for the 2.0 API for tweets

    Note That the Recent Changes of the Twitter Site Seem to Affect Their Response Time to Process Applications for Developer's Account. 
    Due to the Delay on the Twitter Site Response Time, You Need to Apply ASAP to Collect the Twitter Streaming Data in Time


    Twitter Stream Data Collection

    Collect at least 10,000 Tweets Talking about either One of the Following Topics of Your Choice.
    Add More Related Keywords As Needed to Your Chosen Topic to Collect As Many Related Tweets Possible:

    Suggested Topics:

    1. Any New Major Movie or Product That Was Released Recently if Any (For example, IPhone - IPhone 12, IPhone Mini)
    Or
    2. Any Major News (For example, 2024 US Presidential Election)
    Or
    3. President or Any Person of Interest, or Any Two Candidates in an Election
    Or
    4. Covid, Corona Virus, Covid-19 related topics
    Or 5. Any Topics of Your Interest as long as there are big enough to collect more than 10,000 Tweets


    Twitter Data Collection Setting Up: (The Twitter Data Collected will be used for Lab3 and Can Be Used For Your Final Project later)


    NOTE that to collect the Twitter stream data in real time, you need to apply for their developer’s account in the Twitter Deveoper's site
    and get a permission to get credentials for a token and keys. This process usually takes 3-4 days or one week for Twitter to respond.

    You HAVE TO start your application ASAP. Don’t wait until the last day.


    See the example project and sample codes for the step by step procedures for this. Read everything posted in the links below.


    Apply a Twitter Developer's Account ASAP for Twitter Stream Data Collection ! Start ASAP Since it Will Take a Week to Obtain a Twitter Developer's Account !!

    For Your Twitter Developer's Account Application, Choose the most Common Account type. Do not choose an Academic Reserach Account (You Will Be Asked More Questions).

    Do NOT Blindly Copy Those Sample Answers in the Class Webpage for Your Answers ! Rephrase/Modify Them in Your Words For Your Case.

    See Examples of the Answers for the Questions from Twitter to Get a Twitter Developer's Account

    (This is For an Academic Research Account, which You Don't Need to Apply) See Sample Answers for the New Questions from Twitter to Apply Twitter Developer's Account



    For Your Twitter Account Application
    For the Project Site if Asked, Provide Your Lab3 Specification above.
    you Can also Provide the CIS612 Project Site and Research Project Description for Big Data and Data Scientist

    CIS 612 Project Site to Provide in Your Application for Twitter Developer's Account
    To Answer with Sample Research Project Description for Big Data and Data Scientist









    How to Get Twitter Stream in JSON file:

    Step by Step Guide on How to Get Twitter Stream in JSON file in Python (After the Tweepy API Version 4.0 as of Spring 2022) NEW POST !! by TA Yixi Luo
    ****************
    Step by Step Guide on How to Get Twitter Stream in JSON file in Python (Before The Tweepy New Version 4.0)


    How to Collect a Twitter Stream to a JSON File then Insert to MongoDB

    Tutorial on How to Get Twitter Stream to Insert into MongoDB in Python As of 2021 Before the New Tweepy Version 4.0

    Note: This tutorial uses last year (2021)'s Tweepy Streaming Library. It has been upgraded to a new version as of early 2022. See The Tutorial for the New Version of Tweepy Streaming Lib Above
    Blindly copy and paste of the codes in this tutorial wouldn't work because of the new version of Tweepy Streaming Lib


    FAQs for JSON Processing in General





    Other Related Documentations and Examples from Twitter sites and Python for Twitter

  • Twitter site Documentation for Developer
  • How to Get Twitter Stream in Python
  • How to Get a Twitter Stream in Python
  • code examples for collecting Tweets



















    Visualization Tools of GEO Spatial Data -- NOT Required for Everyone !!

    QGIS
    QGIS in Python

    Example to Visualize in QGIS

    Errors: Why Data Points are Mapped into Ocean?

    Example of Lab2 GIS Data Processing in Python: Example Project to Guide How to Process Geo Spatial Data (generated from a machine) for Hot Spot Analysis From CIS 660 Project by Sarvesh Chande
    Lab2 GIS Data Processing Example of NIJ Geo Spatial Data Visulaization Using ArcGIS Map for Hot Spot Analysis NEW POST !
    Example of NIJ Geo Spatial Data Visulalization Using ArcGIS 3D Map for Hot Spot Analysis NEW POST !
    Example of GIS Data Processing in Java Script to Visualize in HTML and NIJ Data file in GeoJSON NEW POST !
    Example of Coordinate Conversion to Oregon North FIPS 3601 NAD83 NEW POST !
    Example of NIJ GIS Data Processing in a Plain Coordinate in Python for Clustering Analysis
    Example of NIJ GIS Data Processing in a Plain Coordinate in R for Clustering Analysis


  • How to Set up Connector between QGIS and MongoDB for GEOJSON Processing From MongoDB in QGIS in Python










  • Other Data Set -- This is NOT for This Semester

    Lab3_3 on GEOJSON Processing: Use GEOJSON data file below instead of Yelp Business Data
    GEOJSON Data File for Lab3_3 on GEOJSON Processing (~ 1GB Zip file)
    GEOJSON Data Description


    See the Class Lecture Notes Section for MongoDB Guides and Set up, and More Details to Learn MongoDB CRUD Operations







































    Lab 4:

    You can choose either Lab4_1 or Lab4_2 below





  • Lab4_1
  • Text Analytics (Text Mining) with Information Retrieval and Natural Language Processing Methods

  • Lab4_1: Text Processing to Build an Inverted Index for Document Vectorization in TF-IDF


  • 1. Building an Simplified Inverted Index in a SQL Server for Lab4_1 (Minimum Required)

    For Inverted Index to Build, You Can Simplify to One Table with (Term, Doc#, TermFreq)

    2. Building a Full Inverted Index Either in a SQL Server or MongoDB as in the Lecture Note as below: (Extra Credit)

    Dictionary table (Term, TotalDocsFreq, TotalCollectionFreq) and Posting Table(Term, Doc#, Term_Freq)

    3. Extra Credit:
    Build Inverted Index with NLP Pipelining for each Sentence to Extract Context Aware Information with POS or/and NER Tagger and Store them either in SQL Server or MongoDB for Retrieval Later



    Input Files for Lab4_1
  • Input File for All Union Addresses for Lab 4_1


  • Example of Inverted Index on State Union Addresses (Note that the Structures of the Index Tables are a little Different in the Example)









  • Lab4_2

  • Lab4_2: Information Extraction using NLP APIs to Create a Knowledgebase in JSON/MongoDB Collection


  • See Extra Credit Lab 4_2 on Information Extraction of Biomedical Document Collection using STANZA NER and OpenIE to Extract TRIPLES to Create a Knowledgebase in JSON/MongoDB Collection


  • Lab 4_2 on Information Extraction from Biomedical Text Documents to Create a Knowledgebase in MongoDB Collection (NOT FOR THIS SEMESTER)



  • Big Data Set and Tools:

    wiki Data set in XML and Html
    Wiki Text Data Sets Complete List and Tools

    Pubmed site XML Description

    32 million Paper Abstracts in XML in Pubmed Baseline






    Text Preprocessing Library in Python SpaCy:

  • Python Text Processing Libs for Text Analysis
  • Liquistic Modules in Python SpaCy
    Lemmatizer in Python SpaCy
    Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy
    How to Code Liquistic Modules like Lemmatizer in Python SpaCy
    Python Example for Basic Text Processing























    (NOT For This Semester !!!!!)
    Lab5

    Sentiment Analysis with Machine Learning for Classification:

    Data Set: Choose a Review Text Data Set Obtained from Social Network sites: Twitter or Yelp Review Data Set


    Classification Goal:

    1. For each review text obtained from Twitter texts, Derive a Preditive Model to Predict (Classify) Whether it is Positive or Negative Review.
    2. For each review text in Yelp Review Data Set, Derive a Preditive Model to Predict (Classify) the scale of the Review in 1 - 5. (star)




    Useful Big Data Analytic Tools

    Choose your System/Tool/Platform to Set Up and Get Used to:

    Python Analytics Tools and Tutorials :

    Machine Learning :

  • Python Scikit Learn
  • Python Scikit Learn for Data Preprocessing
  • Python Numpy Tutorial

  • Other Machine Learning Platforms:

  • Pandas: The Python Machine Learning library
  • Anaconda Machine Learning Platform for Python, R
  • Pandas Python Tutorial Codes
  • Keras: The Python Deep Learning library for Tensorflow, CNTK

  • Basic R Tutorials :
  • R Studio basic Tutorial
  • Examples to start with R Stat Tool
  • Examples of Basic R Stat Tool with Helpful References















  • Lab 5:
  • Lab Assignment 5_1 on Setting Hadoop and Running a MapReduce Job - Word Count
  • Lab Assignment 5_2 on Setting Hive a Parallel Database Server on HDFS and Create DW Tables with Directory Index Structure

  • Lab5 on Configure to Run a Hadoop MapReduce Job and Hive Parallel Database Server
  • For Part 2:
    Data For Hive or Pig Latin: Use either Video Game Sales Data below or Yelp Business.json
  • Global Video Game Sales Data


  • Lab5_1 on Installing Hadoop Updated with Extra Credit Task !

  • For Those who Want to Set Up Your HDFS Cluster On EC2 Amazon Cloud, See the Cloud Section at the end of the Class Lecture Note Section for a Student Account.

    Hadoop Set Up Instruction Sites:
    Hadoop single node setup
    Hadoop Cluster setup
    Map Reduce Tutorial on Hadoop

    Data Set for MR Job:
    NASA HTTP Access Log File
    AWID:Wireless Network Server Log Data Set
    AWID Data Set Avaliable here as well:
    AWID: Wireless Network Server Log Data (around 20 GB zip)

    You can use a Container: Docker (Instead of VM) to Configure a Distributed Cluster for Lab4_1

    Lab Guides: Newest on the Top
    Hadoop Installation on Ubuntu: How to Install Hadoop on Ubuntu (as of 2020)

    Trouble Shooting Tips for Installing Hadoop on VM Do not reformat again to avoid losing name node data
    Hadoop Installation: How to Fix When Data Nodes are not Running (2018)
    Help Site on How to Fix When Data Node are not Running (2018)
    Procedure to How to Execute MapReduce in Eclipse to Run a Wordcount Job (2018)
    Procedure to Install Hadoop and Run a Wordcount Job on Window 2010 (2017)
    For Set up Problems, The following post is helpful.
    Container is running beyond memory limits

    A good Instruction Site for Installing and Running Hadoop
    Installing Hadoop

    If you have a trouble installing Hadoop from the above site with not seeing the name node, You need to delete the temp files created in standalone mode and reformat the namenode.
    For setting up a passwordless SSH, see Ganesh's Lab4 below as well.
    Installing Hadoop and running a wordcount job by Ganesh VAVILAPALLI
    Video for Installing Hadoop shared by Prashant Patel
    Guideline3 for Lab4_1 on Mac (2013)


    MR Programming on Hadoop Lab: (NOT FOR THIS SEMESTER !!!!)
    Lab4_2 on MR Programming on Hadoop:
    Implement Average Temp By Station in MR on Hadoop as in Lab Guide below. Follow the sample codes below.

    Simple MR Hadoop Lab Guides
    MR sample Codes - Average Temp By Zip Code
    MR Sample Codes - Sort By Station ID
    Zip file for station Data Set and
    Zip file for station Data Set and MR Codes







  • Lab Assignment 5_2 on NoSQL on Hadoop / Project Set Up

    NoSQL Systems on Hadoop Lab:

    Part 1: Due By the midnight of the End of the Fourth Weekend of November
    Part 2: Due By the midnight of the End of the First Weekend of December
    All the Extra Credit Labs By the end of the Last week of the class in December!

    You can choose One Parallel NOSQL System for Lab5_2 below.

    Lab5_2 on Using NoSQL Systems: Hive, PIG, MongoDB, HBase, or Spark

    For Part 2:
    Data For Hive or Pig Latin: Use either Video Game Sales Data below or Yelp Business.json
  • Global Video Game Sales Data


  • Example of Creating a Hive Table with Partitioning

    Required Setting for Partition and Step by Step Example for Creating Hive Partitioning Tables with Partition By
    Note that For Creation of Partition Tables, for partition, you have to set this property set hive.exec.dynamic.partition.mode=nonstrict


    Note that Hive has a bug that initialize your Name node whenever you start Hive again. You may lose your data so BACK UP Your Data !
    Most of set up problems are coming from a Version Mismatch between conponents. It may not have documented correctly (remember these are open sources).


    Distributed Parallel NoSQL System Set Up and Basic CRUD Operations in Examples :

    Hive Table Patitions For Data Warehouse
    MongoDB Join in Python Script as Client
    Spark Basic Data Processing
    HBase Join with Hive


    Hive Set Up Guides for Lab4_2 (2019-2020)

    Hive Set Up guide (2020) NEW POST !!
    You need to use new version JDK for a newer HIVE version: oracle – 8 -jdk (istead of using default JDK)

    Common Hive Set Up Probelms and Solutions NEW POST !!

    Hive Installation Guide (2016)



    Data SetS:
    Video Game sales Data for HIVE/PIG

    If You choose MongoDB, Do Aggregate Data Pipeling in Lab3_4 in the Lab3_4 Section Above.
    Data for Mongo DB: Use business.json and review.json from the Yelp site
    JSON Files from Yelp Challenge

    NASA HTTP Access Log File
    AWID:Wireless Network Server Log Data Set
    AWID Data Set Avaliable here as well:
    AWID: Wireless Network Server Log Data (around 20 GB zip)



    NOSQL System Set Up Guides for Lab4_2

    Hive Set Up guide (2020) NEW POST !!
    You need to use new version JDK for a newer HIVE version: oracle – 8 -jdk (istead of using default JDK)

    HBase with Hive Set Up guide

    Hadoop/Cloudera/Hive Installation Guide

    NOSQL System Set Up Guides for Lab5_2

    MongDB/Hive Installation Guide
    MongoDB/Hive Installation Guide, Permission Error Resolution
    MongoDB/Hive Installation Guide
    PIG/MongoDB Set Up Guide



    Spark: Real Time System Set Up Guides:

    Spark Installation Guide with HDFS (When HDFS is already set up on your system)

    Spark Installation/Configuration Guide with HDFS

    Spark Set Up Guide with HDFS

    Spark Installation Guide on Ubuntu (Without HDFS)

    Spark Installation Guide on Win10


    Kafka: Real Time Messaging System Set Up Guides: (2020)

    Kafka Installation Guide on Ubuntu


    Cloudrea: Integrated Real Time Data Analytic Platform: SparkSQL, Spark, Hive on Hadoop
    Cloudera Installation Guide


    NoSql System Set Up Guides for Lab5_2:
    HBase and Cassandra Installation Guide
    HBase on Cloudera Setting up and Tutorial
    Spark Tutorial From Danielle Aring's Report
    VoltDB (In-Memory Relational Database Server) Commands and Stored Procedure Sample Codes


  • Cloud Computing Set Up

    You can create your Big Data Processing Infrastructure on Amazon Cloud for your Project
    Amazon Cloud Account for Students





    Supporting Contents for Labs :

    How to Create a Web Application with Java Based Application Server with MS SQL Server:
  • Set Up Instructions for Java Based Application Server with MS SQL Server
  • How to Create Java Servlets for CRUD Operations
  • How to Create Java Servlets for CRUD Operations


  • Amazon Cloud:
    Amazon RDS (Relational Database Service)
    Amazon Elastic Cloud Computing (EC2) for Web service
    How to Create Amazon Virtual Hosting (EC2) with a Web Server and Amazon Database Server (RDS)

    Microsoft Cloud Azure:
    Trial account for Microsoft AZURE Cloud
    How to Create/Retrieve a Table in Microsoft AZURE Cloud
    Sample Project on Microsoft AZURE Cloud
    How to Create/Retrieve BLOB data in Microsoft AZURE Cloud
    AZURE Cloud
    How to Create a SQL Database Server in Microsoft AZURE Cloud
    Tutorial for MS AZURE Cloud


    Useful Resource Sites:
  • XHTML on WWW.W3.Org
  • XML Version 2.0 on WWW.W3.org
  • DOM XML Parser
  • Microsoft XML DOM Parser Beginner's Guide
  • SAX XML Parser
  • Example Codes of XML Processing with XML Parser (DOM Parser and SAX Parser)
  • XML Editing with OXIGEN
  • Useful XML Resource Sites
  • Beginner's Guide to XML DOM


  • Class Lecture Notes with Tentative Schedule

    Class Chapter / Topic / Specific Objectives / Activities
    1


  • Introduction to Big Data Technologies and Big Data Processing Systems on Cloud

    Introduction to Big Data, Big Data Processing, Big Data Processing Systems, and Big Data Anaytics for AI

    Introduction to Big Data, Big Data Processing, and Big Data Processing Systems for Big Data Analytics

    Overview of Big Data Processing Systems: NO SQL Systems

    Core Builing Phases of AI with Big Data, Big Data Processing and Big Data Analytics


    5-Min Overview on What is Big Data Analytics





    Getting to Start with Examples of Intelligent Systems (AI) with Big Data, Big Data Processing and Big Data Analytics:


    Examples of AI Applications with Big Data

    LectureNotes_2_1: Overview of Question Answering System

    LectureNotes_2_2: Overview of IBM Watson: Question Answering System

    Lecture Notes_2_3: Overview of Sentiment Analysis System


    Examples of Big Data Processing Systems:
    Class Note_2_2: Example of Enterprise Big Data Processing System on HDFS with the Data Warehouse Approach - LinkedIn Data Processing System
    Class Note_2_3: Example of Enterprise Big Data Processing System on HDFS: Yahoo PigLatin




    Examples of Big Data Projects in Big Data Lab:
    Project Example_1: Sharesci - Architecture of an Intelligent Document Search Engine: sharecsi
    Project Example_2: Overview of Intelligent Document Search Engine Using Machine Learning with Natural Language Processing

    Examples of Big Data: 7.2 Million Wiki Webpage Dump in XML format to download to process

    wiki site of sharecsi for Documentation
    github site of an Intelligent Document Search Engine: sharecsi for Source Codes






    For Your Information:

    What are the Most Imprtant Topics to Learn in Big Data ? All of Them Are Covered in CIS593, CIS611, CIS612, and CIS660
    For Comparison: Topics in Big Data Specialization

    Realated Subjects to Be Covered in this Big Data Course
    For Comparison:Topics in Big Data Engineering Specialization






  • 2-5



  • Big Data Processing

    Motivations:

    Examples of Industry Big Data Processing Systems: LinkedIN (and Example of Final Big Data Project)

    Class Note_2_3: Example of Enterprise Big Data Processing System on HDFS: Yahoo PigLatin

    Examples of Data Pipeling for Query Execution Steps in Big Data Processing Systems

    Example of Enterprise Level Big Data Processing Infrastructures and Tasks



    Getting to Start with Examples of Intelligent Systems (AI) with Big Data, Big Data Processing and Big Data Analytics:

    LectureNotes_2_1: Overview of Question Answering System

    LectureNotes_2_2: Overview of IBM Watson: Question Answering System

    Lecture Notes_2_3: Overview of Sentiment Analysis System


    Common Big Data Formats:

  • Structured Data: Data in Table or CSV/TSB Format
  • Semi-Structured Data: Data in HTML, XML, JSON
  • Graph Structured Data: Data in Triples (Nodes and Edges)
  • Unstructured Data: Webpages/Text/Documents/Electronic Books


  • Semi-Structured Data Processing for Big Data Applications (for Serverside Processing)



    Semi-Structured Data Processing Techniques in Web Applications

    Overview of Architecture of a Web Application Platform

    Lecture Note on Semi-Structured Data Model with XML and JSON


    Universal Data Exchange Formats on Internet Between Web Applications As Client and Server -- Platform Independent

    In Semi-Structured Data Model:

  • JSON (JavaScript Object Notation) in Semi-Structured Format
  • XML (eXtensible Markup Language) in Semi-Structured Format
  • HTML/XHTML (Hyper Text Markup Language) in Semi-Structured Format

  • In a Structured Model:

  • CSV (Comma Separate Value)/TSV(Tab Separate Value) -- This is a Structured Format in a Relational (Table) Model


  • Good Example of Web Service Site
    Bad Example of Web Service Site
    Bad Example of Web Application Site




    Common Big Data Formats:

  • Millions of Webpages -- Text Documents with Markup Tags
  • Social Media Server Logging Data -- Twitter or Facebook Server Longging Data -- Millions of Semi-Structured Data in JSON or XML Documents
  • Web Server Logging Data -- Structured or Unstructured Text Documents

    Big Data Examples in Semi-Structured Data in XML or JSON Encoding Format from Real Life Applications:

    Data sets at Yelp Site

    Yelp Data Format: Business.json

    Social Media Twitter Server Generated Data in JSON

    Structure of Twitter Server Logging data in JSON

    Medline Plus for Medical Encyclopedia

    7.2 Million Wiki Page Dump either in HTML or XML

    7.2 Million Wiki Page Dump in XML

    32 million PUBMED biomedical paper collection in XML




    Common Big data Format in Semi-Structured Data Model:

    Big data in Semi-Structured Data Models for Data Exchange Formats between Web Applications in Internet

    Web Application Architecture


    HTML/XHTML/XML Documents (a Collection of Millions of Webpages) as Big Data in Semi-Structured Data Model:

    Markup Languages: HTML, XHTML, XML as in a Semi-Structured Data Model

    Big Data as Semi-Structured Data on Web Application Platforms on WWW


    World Wide Web Consortium (W3C) Reference Sites for MarkUp Languages:

    World Wide Web Consortium (W3C) Standard for Markup Languages: HTML/HTML5/XHTML
    HTML/HTML5
    HTML/HTML5 Syntax by W3 Consortium
    XHTML1 XHTML2 is on the way !



    1. HTML/XHTML:

    REVIEWS on Markup Languages and Basic Web Technologies:

    See See CIS408 Class Lecture Notes for Tutorials on HTML, JavaScript



    HTML
    HTML Basic Summary
    Difference between HTML/HTML5 and XHTML

    HTML References:
    HTML Element List
    HTML Attribute LIST
    Global HTML Attribute LIST
    HTML EVENT Attribute LIST


    CSS:

    Class Note_1_4: Letcture Notes CSS
    Basic Html CSS Tutorial with Examples
    Full CSS Example Site


    URLs:

    Class Note_1_5: Letcture Notes on URLs
    HTML URL Encoding
    URL Query Parameter Example with Form Element
    How Your Personal Webpage Works with the Webserver on the EECS Host in URL
    HTML File Path Rules in URL


    Understanding Form Element with Http Get method to send parameter values from Client site html page (by Web Browser) to a Serverside program (with Web Server)

    Form object and Submit Type with get Method



    Examples of Big Data in Webpages as Semi-Structured Data Model:

    Webpage Processing Techniques (for a Large Collection of HTML source files or in XML data files) with Semi-Structured Data Model:

    Examples of Big data in Web data (either in a large collection of html source files or in XML data files) to process to build an Intelligent Queston Answering System:

    Information Service Site for the US President State Union Addresses in HTML
    Trip Advisor Site with Hotel Reviews in HTML
    Example of a Wiki Webpage
    an Example of Wiki Webpages in a HTML document
    7.2 Million Wiki Page Dump either in HTML or XML

    Bad Example of Web Service Site to Collect Data from







    Big Data Processing Technologies for a Large Collection of HTML Pages as Semi-Structured Model


    Lectures:

    Semi-Structured Data Model:
    Class Note_10: Lecture Notes On Semi-structured Data Model and XML

    Document Object Model (DOM):

    HTML/XHTML/XML in Semi-Structured Data Model - Document Object Model (DOM):


    Running Example of Webpage Processing Method to Extract Information from Large Corpus of Webpage Documents:

    The websites for the collection of the US President State Union Addresses
    Example of an HTML Document as Source Text File to Process



    Document Object Model (DOM) for Semi Structured Data Processing

    DOM is the Implementation of Semi-Structure Data Model for HTML/XHTML/XML


    Big Data Processing Techniques for a Large Collection of HTML Documents or XML Documents

    HTML/XHTML/XML in Document Object Model (DOM)

    DOM Lectures:

  • HTML in Document Object Model (DOM):

  • Class Note_3: Intro to Document Object Model (DOM) for HTML and JavaScript
    Class Note_3_1: Example Picture of DOM Tree of Web Browser for HTML and JavaScript
    Example of Document Object Model (DOM) for HTML and JavaScript

    Node Types of DOM Tree: (From W3C Schools Tutorial on HTML DOM)

    DOM Document Node Type for Document Element HTML
    DOM Element Node Types and Methods/Properties
    DOM Attribute Node Type and Properties
    DOM Text Node Type

    HTML DOM Reference List in JavaScript:
    HTML DOM Document and Method List
    HTML DOM Element and Method List
    HTML DOM Attribute and Method LIST
    Important Note:
    DOM Query Result in Collection vs List


    MDN: HTML DOM Intro
    MDN: DOM Parser for HTML/XML
    MDN: Useful Example Sites of DOM Objects
    MDN: HTML DOM API


    Official DOM References on W3C:
    Official Document Object Model (DOM) Specification for HTML/XHTML/XML by W3 Consortium


    IMPORTNT NOTES !!!

    Note that the code examples in JavaScript in this section Are Only for Client-side Processing for Dynamic Webpage Changes Where a Web Browser Automatically Parses and Builds a DOM tree for the Html file to Execute.
    This client side Javascript examples are only for your understanding.
    All the labs and Final Project in this course require Server-side Data Processing in any modern Server-side (script) languages such as Python, C++/C#, PHP, Java, or Node JS (Server-side JavaScript with File I/O)

    Code Examples of DOM Tree :

    Example of DOM property InnerHtml
    Example of How to Find DOM Objects: getElementByTagName
    Example of How to Access DOM objects and values: childen element count
    Example of HTML DOM Object with button
    Element.className

    Example of How to Access DOM Atrributes: namednodemap: getnameditem
    Examples of HTML DOM Object creation
    Examples of HTML DOM Object Create and Append
    Example of How to Create Text Node
    How to create DOM Objects and Replace: replaceChild


    Javascript const and let: Three ways to declare a variable and Differences between let and const


    How to Access an attribute value of Form Element
    How to Create Form object






    Difference between children and childNodes (as of 2022):

    Difference between DOM Element children (Element children only) and childNodes (All the Child Nodes -- Element Nodes, Text Nodes, Comment Nodes)

    DOM Element children in Collection:
    The Children method returns a Collection (in Array) of Elelement Child Nodes (and its subtree).
    Exampe: DOM Element children (Element children only)
    Example: DOM Element children (Element children only)


    childNodes in NodeList:
    The childNodes returns a NodeList (in Array) of all the nodes (Element and Text nodes).
    Whitespace inside elements is considered text nodes. In the example below, index 0, 2 and 4 in the Node List of childNodes of myDIV are text nodes.
    Difference between DOM Element children (Element children only) and childNodes (All the Child Nodes -- Element Nodes, Text Nodes, Comment Nodes)
    Example of DOM Element childNodes (All the Child Nodes -- Element Nodes, Text Nodes, Comment Nodes)
    Example of DOM Element Select with childNodes (All the Child Nodes -- Element Nodes, Text Nodes, Comment Nodes starting from an empty Text Node as index 0)




    HTTP:

    Lectures:

    HTTP: How to Communicate between a Web Browser and a Web Server:

    Class Note_17: Introduction to HTTP Protocol on How a Web Browser Brings(Requests) Source Files (in URL) in HTML and JavaScript from a Web Server
    Lecture Note_5_1: more on HTTP

    This a Form element Example and an Example of an AJAX request to be transformed to a http request by the web browser.
    Example: From Form Element to HTTP Request by a Web Browser





    How to write HTTP (in your JavaScript with HTML) to Communicates (Request Files as a Client) with a Webserver

    AJAX (Asynchronous JavaScript and XML):


    Lectures:

    Class Note_18: Controller as Server Communication with Client in JavaScript Using AJAX

    AJAX Intro
    AJAX XMLHttpRequest Object with Methods and Properties
    Class Note_11_2: AJAX Intro and Syntax Quick Summary


    AJAX Code Examples:

    A form element like this is transformed to an AJAX request like this by a web browser.

    Client Request:

    AJAX Send a Request to a Sever with GET/POST methods
    AJAX Examples - XMLHttpRequest
    URL Query Parameter Example: GET method
    AJAX Examples - POST method

    Server Response:

    AJAX Server Response (Scoll Down to Look at the Example for Using a Callback Function)
    AJAX Examples - Callback Function
    AJAX Example - retrieving Server Response
    AJAX Examples - Processing XML Response
    Example with a Callback Function for ResponseText Handling
    AJAX Examples - Accessing Server Response Header


    Basic XML Data Processing Examples in JavaScript with AJAX in HTTP in a Client Site:
    XML File Example: cd_catalog.xml
    AJAX XML File Processing Examples
    AJAX Examples - XML File Processing
    AJAX XML Processing Example


    AJAX and JSON Data Processing:
    AJAX and JSON Data Processing
    AJAX and JSON Array Processing






    IMPORTNT NOTES !!!

    All the labs and Final Project in this class require Server-side big data processing in any modern Serverside (script) languages such as Python, C++/C#, PHP, Java, or Node JS (Server-side JavaScript)

    When You Build a Server-side Application with any Programming Language to Process HTML/XML Documents with DOM and XPATH for an Big Data Application,
    You Have to Call a DOM Parser to Build a DOM Tree for Each HTML/XML Document !

    In any other seperate script/programming lanaguages like Python, Java, or Node JS, you have to use(call) a DOM Parser that will build a DOM Tree from your input document like html or XML file.



    Useful DOM Tree Query Features for Big Data Proceesing:

    XPATH: See the XML Section Below for the XPath Lectures !
    Setting Up lxml in Python for DOM with XPath and Code Examples ***************
    XPath in JavaScript

    Query Selector with CSS, JS RegExp:
    Document.querySelector with CSS Selector or Rex
    get elements by selector: querySelector with CSS Selector in Python

    Review on CSS and JS RehExp:
    HTML CSS Rules to Select Elements
    Letcture Notes JavaScript Basics (See Page 18 - 20 on Regular Expressions)
    JS RegExp
    RegExp: ignoreCase property of i modifier


    DOM with XPATH Set Up in Any Script/Programming Languages for Server Side Applications

    Note that All the RECENT Web Browsers Have DOM and XPATH Features in their Debugger by Default.
    You don't Need to Do Any Extra Set Up to Use XPATH in your JavaScript in a Client Code for HTML Processing in your Webbrowser.

    Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use !
    For the Labs to Write a Server Application in this course, You Need To Add DOM Parser Related APIs for Your Scripts/Programming Languages to Do:

    1. Call a DOM Parser to Build a DOM Tree
    and
    2. Use XPATH Query Methods to Retrieve DOM Objects to Extract Information (Tag Names or Text) in the Application Codes as in the Examples below.


    Useful Webpage Processing Tools with Built-In DOM Parser and XPath Query Feature:

    Setting Up lxml in Python with DOM Parser and XPath
    Beautiful Soup Documentation for Document Processing for HTML/XHTML, XML in Python
    DOM Example: Webscrapping with Python BeautifulSoup

    Selenium-Python: Setting Up DOM with XPath for Python
    Selenium-Python: XPath API to Locate Elements in Python


    Code Examples of DOM and XPath for Document Processing in Python:

    Python XPath Code Examples
    Webscapping with DOM, XPath for Python with Beautiful Soup (From the Stanford Class)


    JDBC (Java Database Connectivity)/ ODBC (Open Database Connectivity):********************************

    Lecture Notes for MySql with JDBC:
    MySQL Applications Using Java & JDBC
    Database Programming with JDBC and Java

    Tutorial for MySql with ODBC:*************************************
    Architecture of Database Applications with ODBC
    pyobdc: How to Connect MySql Database with pyodbc in Python


    ODBC(Open Database Connectivity) for Communication Between PHP Server and MySql Database Server:

    List of All the ODBC MySqli Classes and Methods with Code Examples for PHP Server and MySQL Server Communication ***

    PHP Basic Example Codes Using ODBC(Open Database Connectivity) mysqli to Connect to MySql ***

    ODBC MySqli: Connection class ***

    ODBC MySqli: Statement class ***

    ODBC MySqli: Prepared Statement class ***
    ODBC MySqli: Prepared Statement with parameter binding ***
    ODBC MySqli: Prepared Statement and Accessing the SQL Result of Prepared Statement ***












    Think about What Would Be a Better Way to Send the Contents of the Webpages of the Collection of State Union Addresses to Application Servers ???


  • XML (eXtensible Markup Language):
  • XML as an Universal Data Exchange Format in Semi-Structured Model

    XML is an Universal Data Encoding Protocol to Seperate Data/Database (Data and their Scheme) from Presentation (Displaying instruction Codes) in Html

    Example of Real-world data at NIH in XML:
    32 million PUBMED biomedical paper collection in XML

    Industry Example of XML data files as Big data (a big XML data file to encode a database) to process to build an Intelligent Queston Answering System:
    7.2 Million Wiki Webpage Dump in XML format to download for a Big data Project


    XML (eXtensible Markup Language) Documents in Semi-Structured Data Model

  • XML in Document Object Model (DOM)

  • Lectures:

    Semi-Structured Data Model for Universal Data Exchange Formats

    Lecture Note on Semi-Structured Data Model with XML and JSON

    Better Way to Send the Contents of the Webpages of the Collection of State Union Addresses to Application Servers !


    Class Lecture Note: Introduction to XML

    Well-Formed XML Syntax

    Tutorial for XML with Examples of XML Data
    simple.xml file with XSLT file: Simple Example of XML Data file with XSLT Instruction file
    simple.xsl file: Simple Example of XSLT Instruction code file for XML Data

    Tutorial: How to Use XML
    Tutorial: XSLT (XML Stylesheet Language) Transformation
    Example of Data in XML and Presentation Instructions in XSLT
    XML DOM Tree Example
    Data Processing for XML DOM




    1. Information Extraction Techniques from a Large Collection of XML Documents as Semi-Structured Data in DOM

    On Recieving XML Data in a Text Format from a Server site, Client Site Has to Convert a XML Data to Objects in DOM Tree Using XML Parser !!
    XML Parser to build DOM tree from xml data (file) is available in any Langauge (Python, C#, Java, PHP for example) to write a Server
    Code example:
    XML Parser in JavaScript to Convert a XML Data to a DOM When a Client Recieved a XML Data in a Text Format in Javascript for Client


    Data Processing for XML DOM
    Example of XML Document Object Model (DOM) Tree for XML Data







    Information Extraction Methods for Big Data Applications

    How to Select(Query) XML DOM Elements

    Useful DOM Tree Query Features:

    XPATH:

    Lecture Notes On XPATH

    Tutorial:
    XPATH To Query (Serach) XML DOM Tree
    XPATH Syntax

    Tutorial: Summary on XPath with Examples
    W3 Tutorial: XPath Quick Syntax Summary


    W3 Consortium on XML with DOM and XPath References:
    Complete List of DOM with XPath in www.w3.org


    Important Notes on Information Extraction from HTML or XML Documents for Labs:

    When You Build an Application Server to Process HTML/XML Documents in DOM with XPATH in a Server-side (Script) Language, You Have to Call a DOM Parser to Build a DOM Tree for Each HTML/XML Document !

    In any other seperate script/programming lanaguages like Python, Java, or Node JS, you have to use(call) a DOM Parser that will build a DOM Tree from your input document like html or XML file.

    Note that You have to Use a DOM Parser to Build a DOM Tree from your Input HTML or XML file to be able to use XPATH query !!
    In any script/programming lanaguages like Python, Java, you have to use(call) a DOM Parser that will build a DOM Tree from your input document like html or XML file.


    DOM Parser with All the DOM Methods and Query Features/APIs like XPATH is available in any modern programming languages - Python, Java, C++/C# for Big Data Applications


    XPath API is available for Querying HTML DOM Tree as well as XML DOM Tree either in Web Browser Debugging Tool or any Programming Languages - Python, Javascript, Java, C#, PHP:

    Example Code for Setting Up DOM with XPath in Python ***************************
    Tutorial: XPath Syntax ***********************


    XPATH Set Up in Any Script/Programming Languages for Server Side Applications

    Note that All the RECENT Web Browsers Have XPATH Features in their Debugger by Default.
    You don't Need to Do Any Extra Set Up to Use XPATH in your JavaScript in a Client Code for HTML Processing in your Webbrowser.

    The Labs in This Course is to Write Tasks to Be Done in an Application Server, You Need To Add DOM Parser Related APIs for Your Scripts/Programming Languages
    Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use !

    To Do 1. Call a DOM Parser to Build a DOM Tree for Each Document in Your Data Collection
    and
    2. Use XPATH Query Methods to Retrieve DOM Objects to Extract Information (Tag Names or Text) in the Application Codes as in the Examples below.





    XML DOM Processng Examples in Different Server side Programming Languages:

    Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use !


    XML API for Server side Applications: DOM Parsers to Download for Programming in JAVA:

    Example Codes of XML Processing with XML Parser in Java

    Example Codes of Mobile APP On How DOM Parser and SAX Parser Work to Process XML data in JAVA


    XML API for Server side Applications: DOM Parsers to Download for Programming in JAVA:

    SAX XML Parser
    DOM XML Parser
    Which DOM Parser jar file Should I Download
    MDN: DOM Parser for HTML/XML


    Code Examples of Conversion from XML to Table in PHP





    Useful DOM Parser and XPATH APIs with Example Codes in Python:

    Setting Up DOM with XPath with lxml in Python ******************************
    BeautifulSoup Documentation for HTML, XML in Python
    Python XPath Code Examples
    Webscapping with DOM, XPath for Python with BeautifulSoup (From the Stanford Class)

    Selenium-Python: Setting Up DOM with XPath in Selenium-Python
    Selenium-Python: XPath API to Locate Elements




    Example Codes of DOM and XPath in JavaScript:
    Note that Although the code examples here are in JavaScript, but All the DOM API and XPATH Query Features are available in any modern script languages - Python, Java, C++/C#

    DOM with XPath Setup Guide in JavaScript with Node JS
    List of XPath Methods in JavaScript with Node JS
    MDN Site for DOM with XPath in JavaScript
    MDN Site for Code examples of DOM XPATH in JavaScript


    MDN: DOM Parser
    MDN Site for Code examples of DOM in JavaScript
    MDN: How to Use XPath in JavaScript





    3. Features of Relational Database Server for Big Data Processing

    Automatic Table Creation in a Database Server:***********************************

    pyodbc(Open Database Connectivity) in Python to Connect MS SQL Server :

    Installation pyodbc to Connect to MS SQL Server in Python
    pyodbc to Connect to MS SQL Server in Python
    How to Connect to MS SQL Server in Python
    pyodbc: Examples of ODBC in Python

    How to Create a Table in a SQL Server from CSV/TSV file*************************************

    How to Create a Table with Bulk Insert with MS SQL Server
    Example of a Script to Create Multiple Tables with Bulk Insert in MS SQL Server
    Example of how to import csv file to a MySql table in MySql Server
    Example of how to export table to csv in MySql Server





    Earlier Features to Handle Big Data in Relational Database Server

    Data Types in SQL Server: BLOB (Binary Large Object) or Text/CLOB (Character Large Object) in MySQL or MS SQL Server

    How to Create and Insert to a Table with a Column of Large Text Data or Image Data in a Relational Database Server:

    Difference between blob and clob datatypes
    What is TEXT data type in MySQL
    Text Data Type in MySQL
    how to insert blob from a file in mysql using-loadfile-method
    insert-retrieve-file-image-as-a-blob-in-mysql



    Industry Example of Database (Semi Structured in Table format) of Extracted Information :

    Google Big Table for Big data processing
















    XQUERY:
    Semi-structured XML Data Query Language
    Lecture Notes On XQuery




    Object Relation Mapping (ORM):

    ORM Mapping from Semi-Structured Data to Structured Relational Database Scheme:

    Transformation of Semi-Structured Model to a Relational Model in the Third Normal Form



    Reviw:
    Realtional Model and Constraints

    The First Normal Form Rule in Relational (Database) Model

    Design Rules of Relational (Database) Model


























    4. JSON (JavaScript Object Notation) as Semi-Structured Model:

    Universal Data Exchange Format: JSON in Semi-Structured Model

    JSON (JavaScript Object Notation): http://www.json.org/

    Big Data in JSON Format in Real-life Industry Applications:

    Data sets at Yelp Site
    Yelp Data Format: Business.json
    Structure of Twitter Server Logging data in JSON


    Big Data in Semi-Structured Data Model in JSON

    JSON Lectures:

    Class Note_10: Lecture Notes On Semi-structured Data Model and Comparisons to HTML, XML, JSON, Relational Table

    Class Note_11: Quick Tutorial on How To Process JSON Data


    JSON Syntax:
    What is JSON ? How to Exchange and Process JSON Data
    JSON Syntax vs JavaScript Object Syntax for values
    Example of JSON Data: One JSON Object


    Example of json data format of a real life application : Yelp One Business Data Format


    How to Exchange JSON Data for Big Data Processing:

    1. Parse: On Recieving JSON data (.json file): Parse JSON Text data to Convert to ==> JSON Objects in Applications in any language to Extract infomation (scheme and values) to Create a Database

    2. Stringify: To Send Database infomation: Encode(Build) JSON format in JSON objects from Database then Stringify to Convert JSON Objects to ==> JSON Text data (.json file) to Send


    How to Exchange and Process JSON Data: JSON.parse and JSON.stringify
    JSON.parse
    Processing Date in JSON
    JSON Data Processing -- JSON.stringify

    JSON Object
    Accessing Properties(Names) of JSON Object
    Accessing Property Values of JSON Object
    Accesing and Changing JSON Object
    Delete Object in JSON
    JSON Array - Nested Array


    JSON Processing Examples:

    Client side in a HTML with Javascript
    JSON Data Processing with HTML in a Client Side

    Handling in Server side application codes:

    Python:

    1. To Parse JSON data to Convert to ==> Python objects: json.loads()
    2. To Stringify Python objects to Convert to ==> JSON text: json.dumps()

    Intro to JSON Data Processing in Python
    Example of JSON Data Processing in Python

    JSON Data Processing with PHP in Server Side
    Lecture Note6_5_5: On How to Process JSON data in a Mobile App






    Object Relation Mapping (ORM):

    ORM Mapping from Semi-Structured Data to Structured Relational Database Scheme:

    Transformation of Semi-Structured Model to a Relational Model in the Third Normal Form






















    7-8




  • Big Data and NoSQL Big Data Systems:MongoDB



    Semi-Structured NOSQL Database Management Systems for Big Data Processing

    MongoDB:

    What is MongoDB?: Overview of MongoDB Atlas (Cloud instance), MongoDB Community (the Simplest Version), and MongoDB Enterprise (Stand alone)

    Mongo DB Setup Related:

    MongoDB Installation:
    MongoDB Download and Installation
    MongoDB Enterprise Download
    MongoDB Installation Guides on Window
    Or
    MongoDB Community Version Download and Installation


    MongoDB Clients Download - MongoDB Shell (Command-line) or MongoDB Compass (GUI)




    MongoDB Manual How to Set up Client and How to Use - Getting Started:
    Full Menu of MongoDB Manuals for Getting Started
    MongoDB Shell (mongosh)
    Download and Install MongoDB Client: MongoDB Shell (mongosh)
    MongoDB Shell (mongosh): Download and How to Connect to MongoDB from the Client mongosh
    Start your client: MongoDB Shell (mongosh) to Connect to MongoDB Server
    Old MongoDB Shell (mongo) Replaced by mongosh

    How to Install pymongo driver and Connect to MongoDB Server in Python application as a client


    How to Import JSON Data File to MongoDB :

    How to Import a json file to MongoDB either in Command line Client or in Python or Java

    How to Import Data file to MongoDB
    Mongo import

    Platform Specific Set up Guides
    Node JS with Mongo DB Setup Guide
    Node JS with PostgreSQL Setup Guide
    Web Application Using Node JS with MS SQL Server and Mongo DB (From a Master Project: Building a Recommendation System usng Amazon and Walmart Product and Review Data Sets)

    Complete Mongo DB Documentation Site






    Lectures:

    Lecture Notes on Overview of MongoDB
    Mongo DB Comparison to Sql Terms and Queries **************************



    Review of Basic SQL Processing on Relational Database:

    Lecture Notes on SQL Processing Semantics ************************

    Definitions of SQL Operators to Process SQL: Selection (Filtering), Join, Union **********************





    Mongo DB Intro with Examples:
    Introduction to Mongo DB
    Mongo DB: Databases/Collections
    Mongo DB: Document Format
    Mongo DB: CRUD Operations ************************
    Mongo DB Getting Started
    Mongo DB: How to Create a database



    Mongo DB Query Compared to SQL:
    Mongo DB Mapping Chart to Compare to SQL *************************

    Mongo DB Basic CRUD Operations:
    Mongo DB: CRUD Operations
    MongoDB Operations: InsertMany
    MongoDB Operations: UpdateMany
    MongoDB Operations: UpdateMany with Aggregation Pipelining
    Mongo DB Operations: ReplaceOne

    Mongo DB Query: Select Operation
    Mongo DB Select Operation on embedded documents *****************************
    Mongo DB Select Operation on arrays *********************************
    All and ElemMatch Operators on array of embedded documents *****************************
    Mongo DB Select Operation on array of embedded documents *******************************


    Example Runs of Mongo DB Queries *************************************


    MongoDB Aggregation with Pipelining:

    MongoDB Aggregation Mapping Chart to SQL*********************************
    Mongo DB: Aggregation Pipelining **************************************
    Mongo DB: updateMany with Aggregation pipelining to add a new name and value into the existing documents (Scroll down to see the examples)
    Mongo DB Operator $unwind to unpack an array list*************************
    Mongo DB Join Operator: $lookup ***********************************
    Mongo DB Join Operator: $lookup with $unwind for array *******************************


    Tutorials in Python:

    How to Use MongoDB in Python for CRUDE: DB/Collection Creation, Insert, and Querying

    How to Import a JSON file to MongoDB in Python

    Example of MongoDB Aggregation Pipeling for Word Count

    How to Save MongoDB Query Results

    How to Save MongoDB Query Results into a variable


    Examples:
    Example of MongoDB Aggregation Pipeling for Word Count

    Application Code Examples in Python with MongoDB and SQL Server

    Example Codes of JSON Data Processing in Python with MongoDB


    Mongo DB Resources:
    Class Note_22_3: MongoDB Resources
    Mongo DB Documentation




    Semistructured Database System - MongoDB: More Advanced


    MongoDB Selection Query:
    Mongo DB Basic Selection Query: $find operator with Condition expressions
    Mongo DB Query (select): $find Example Explained

    MongoDB Query with Comparisons to SQLs:
    Mongo DB: Query find operator basic
    Mongo DB: CRUD Operations
    Mongo DB to SQL Basic Query Mapping Chart (Comparisons to SQL in Examples)

    MongoDB Expressions for Conditions:
    Mongo DB Condition Using $expr With Conditional Statements
    Mongo DB Logical Operators
    Mongo DB Condition with Logical Operator $and

    MongoDB Aggregate Pipelining:
    MongoDB: Intro to Aggregation Pipelining
    MongoDB Query/Aggregation Pipelining Operator Mapping Chart to SQL ************************
    Class Note_24_1: Summary of MongoDB Query/Aggregation Ppelining Exampes with Comparison to SQL **************************
    Complete List of MongoDB Aggregate Operators for Pipelining

    MongoDB lookup Operator for Join with Aggregation Piplelining :
    MongoDB $lookup operator for Join: Join returns null matching as well (Left Outer Join Semantics)
    MongoDB $lookup operator with $unwind for Array Valued Field as Join column
    MongoDB $lookup operator with $let for Multiple Fields for Join Columns
    MongoDB let operator

    MongoDB CRUD -- More Advanced:
    MongoDB UpdateMany Operation with Aggregation Pipelining and $set as $addfields, $unset Operators
    MongoDB Add Field with Aggregate operation Pipelining
    MongoDB Add Field with Aggregate operation Pipelining
    MongoDB Add Field with Aggregate operation Pipelining


    Examples of Big Data Project with MongoDB and Yelp Data

    Project Presenation
    Code Examples with MongoDB and Angular JS


    Examples of Big Data Applications from a Previous Research Team in Big Data Anaytics Lab

    sharesci: Intelligent Search Engine in Natural Language

    sharesci: Content Based Document Search Engine on wiki pages and research papers
    sharesci: Content Based Document Search Engine with Machine Learning

    sharesci github site for Source Codes
    sharesci wiki site in bigdata lab for documentation
    sharesci Project Report for Source Codes


  • 10-11



  • Unstructured Text Data Processing


    Unstructured Text data Processing techniques for Information Extraction (IE) Methods

    Information Extraction Techniques for Unstructured Text

    Unsupervised Information Extraction Methods

    Lectures:

  • Semantic Web:RDF (Resource Definition Framework) to Represent Knowledge in Web Page Contents

  • Graph to Represent Knowledge in Web Page Contents in TRIPLES (Subject, Predicate, Object):

    Semantic Web in RDF to to Represent Knowledge in Web Page Contents


    More on RDF Later in the Knowledge Representation Section Below!!! *************************************









    Big Data Analytics: Analysis of Unstructured Text


    Motivation:

    Applications of AI:

    LectureNotes: Overview of IR (Information Retrieval) based Question Answering System: Google Search Engine (AI that Can Answer to a Question (Query) in a Natural Language)

    LectureNotes_2_2: Overview of Knowledge-base Based Question Answering System (IBM Watson)

    Lecture Notes: Overview of Sentiment Analysis System




    Lectures on Unstructured Text Data Processing:****************************************

    Lexicon Based Approach with Document(Text) as Bag of Words Model:

  • Lecture Notes on Information Retrieval with TF-IDF for Text Document/Webpage Analysis ******************


  • Similarity Measures:

    Jaccard Similarity with Example ***********

    Cosine Similarity with Example ***********

    Similarity/Disimilarity Measures|




  • Inverted Index: Efficient Index for Unstructured Text Processing

  • Lecture Notes On Inverted Index in Google (Stanford)

    Example of Inverted Index Scheme



    Google's Inverted Index on Their Big Data - Entire World Wide Webpages:

    Google NGram Viewer : Inverted Index for all N gram words

    Google NGram Viewer : Inverted Index for all N gram words

    How to Use Google NGram Viewe



    Example of Inverted Index Tables for Text Analysis in State of Union Addresses










    Basic Text Preprocessing with Natural Language Processing (NLP) APIs for Implementation of Text Analysis Applications:


    Introduction to Text Processing

    Basic Text Preprocessing with Natural Language Processing (NLP) Methods



    Text Cleaning for Preprocessing:
  • Common NLP Preprocessing Tasks to Be Done for Text Analysis********************


  • Stemming/Lemmatization:

  • Common NLP Preprocessing Task: Stemming-Lemmatization in NLTK python
  • Common NLP Preprocessing Task: Stemming-Lemmatization from Stanford NLP Group



  • Example Project/Application on Twitter Text Mining and Sentiment Analysis in R




  • IMPORTANT NOTES !!!***************************************
    Note that You Have to Decide Which Text Cleaning Preprocessing Should Be Applied or Not Applied Depending on Type of Text Analysis

  • For TF-IDF Based/Lexicon Based Text Vectorization, You Should Apply All the Text Cleaning Preprocessing, Stemming/Lemmarization, and Stopword Removal.

  • For Context Aware NLP Methods Such as POS Tagging, or Word2Vec Training for Word Eembedding, You Should Not Remove Each Sentence Strcuture, So Should Not Apply These Common NLP Cleaning/Preprocessing -- Stopword Removal, Stemming/Lemmarization Which Remove Structures of Sentences.






  • To Get Lexicon Word(Dictionary) Lists
    WordNet -- Largest Lexical Database of English
    WordNet and Synset Download


    General (Short) NLP English Word List for Simple Text Analysis Applications:

    Stop word List
    Positive word List
    Negative word List





    Sample Big Data Projects with Text Analysis:

    Sample Big Data Project: Building a Content Based Document(wiki Pages) Search Engine System
    Sample Big Data Analytic Project: Building a Intelligent Document Ranking System Using Machine Learning
    Sharesci Github Site for Code Examples
    Sharesci Wiki Site for Documentation




    Advanced Study on Clustering Analysis Using Similarity Measures for Document Categorization

    Clustering Analysis Concept
    Clustering Algorithms in scikit-learn







    5-6





    Problems (Limitations) with Lexicon Based TF-IDF for Text Analysis:

    Phrase (N-Gram Word) Identification
    Negation Handling
    Can NOT Identify Relationships among Terms - Synonyms (Similar Meaning) or Antonyms (Opposite Meaning) of Terms
    Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence - Polysemy Problem !
    Can NOT Identify New Terms, Common Slangs or Changining Relationship Between New Terms (ex: Data Analytics) and Old Terms (ex:Data Mining)
    Can NOT Identify Sarcastic Contexts
    Can NOT Identufy the Order of Words in a Sentence !!!!




    Some Solutions for the Problem that Lexicon Based TF-IDF Methods Can NOT Identify Different Meanings of a Same Term by the Different Context of a Sentence

    Context Aware Natural Language Processing(NLP) Methods:

    Information Extraction Methods with Natural Language Processing(NLP) Technologies for Unstructured Texts
    Context Aware (Meaning of a Word from a Sentence Context) NLP Methods for Information Extraction:


    Context Aware Learning with Sentence/Paragraph/Document as it is

    Context Aware Natural Language Processing(NLP) Methods:

    Lectures:

  • POS (Part of Speech) Tagging

  • Lecture Notes on POS (Part of Speech) Tagging (Stanford)
    English Grammar
    UPenn Treebank POS Tagging List






  • NER (Named Entity) Tagging

  • Lecture Notes on NER (Named Entity Recognizer) (Stanford)
    Software NER Tagging Parser:

    Stanford NLP NER (Named Entity) Tagging

    Stanford NLP NER Time Tagging - SuTime



    Stanza: Most Recent Stanford NER

    Stanza: Stanford NER on Biomedical Domain



    Useful Natural Language Processing (NLP) APIs:

    Core NLP - Stanford NLP Description with Example **************

    Core NLP online Run Test Site (Stanford) *******************

    Stanford Core NLP Data Pipelining for Information Extraction**************








    Important Big Data Text Processing Techniques: Information Extraction(IE) Methods:



  • Open IE (Information Extraction) Parser

  • OpenIE Example from Stanford NLP Research site
    Stanford NLP Open IE (Information Extraction) for Triple Extraction






    Tutorial with Stanford Core NLP Pipelining

    Stanford Core NLP Data Pipelining for Information Extraction

    Stanford Core NLP Pipeline Concept for POS and NER tagging This site is also Github repo site as well

    Stanford Core NLP Demo: POS and NER tagging Examples

    Stanford Core NLP Run Site







    Python NLP Text Processing APIs to Build Data Pipeling for Text Preprocessing for Information Extraction :

    Python spaCy Libs for stemming, lemmatization or NLP Preprocessing Methods below !

    Introduction to Text Processing in Python NLP Lib spaCy:

    Use Python NLP Lib called SpaCy to build a data pipelining with NLP APIs for Text Analysis

  • Download and Installing spaCy
  • Introduction to spaCy101 for Text Proceesing with NLP Methods
  • Python spaCy Objects for Text Processing in Data Pipelines
  • spaCy: Building Data Pipleines for Text Proceesing Tasks with NLP Methods
  • spaCy NLP Features
  • spaCy NLP Method: Rule Based Matching


  • Tutorials of spaCy for Text Proceesing with NLP Methods



  • Other Useful Core NLP Software to Download to Use

  • NLTK (Natural Language Tool Kit)
  • Stanford Core NLP APIs (from Stanford NLP Research Group)






  • WordNet Site (by Princeton) to download NLP Software:

    WordNet by Princeton
    Download WordNet for Sentiment Analysis on Aspect/Feature Opinion Analysis

    WordNet Sinset Similarites for All POS Tagging for Dictionary Construction for Text Analysis











    How to Build a Sentiment/Opinion Analysis Application with Text Analysis Techniques -- For Final Project !!
    Lecture Notes_2_3: Overview of Sentiment Analysis System


    Overview of Text Processing Tasks for Sentiment Analysis of User Review Texts :

    Fundamental Papers/Lecture Notes to Learn Basic Methods for Sentiment/Opinion Analysis:

    Lecture Notes On Summarizing Frequent Feature(Aspect) words and Opinions in Review texts
    Paper: Summarizing Features of Product Reviews for Sentiment Analysis

    Lecture Notes On Aspect Analysis Approaches




    Example Big data Project I:

    Pang's 2002 Paper: Sentiment Analysis on IMDB : Example Project
    Presentation: Sentiment Analysis on IMDB (From Pang's 2002 Papers)


    Text Processing Tasks for Product Review Analysis:

    How to Do Sentiment Analysis from my text Corpus

    Project Examples of Vectorization of each document (to Label Ground Truth of a Training set for a Machine Learning Classifier):

    Project Example I: Sentiment Analysis of Yelp Business Review with Machine Learning from a Research Paper
    Project Example II: Sentiment Analysis of Yelp Business Review with Machine Learning
    Project Example on LDA for Topic Discovery and Word2Vec for Similar Words with Trip Advisor Hotel Review Data












    Examples of Text Analysis Tasks for Big Data Applications (Intelligent Systems):

    Example of BETTER Big Data Project II:

    For Semi-Supervised Learning

    Well Known Ranking Algorithms to Score a Sentiment Score of Each Token Word for Vectorization of Document for Labeling (Ground Truth) of a Training Set for Sentiment Analysis:

    1. Senti-WordNet for General Context of Texts:

    Senti-WordNet for Sentiment Analysis
    Presentation of Paper: Senti-WordNet for Sentiment Analysis
    Paper: Senti-WordNet for Sentiment Analysis





    2. Social Media (Twitter) Text Specific Ranking Algorithm:

    Ranking Algorithm: VADER for Sentiment Analysis
    Compound Score of VADER Ranking Algorithm for Sentiment Analysis
    Paper: VADER for Sentiment Analysis






    Example of Semi-Supervised Learning Based Sentiment Analysis Proejct:

    2020 Presidential Election Prediction from Social Network Twitter Analysis

















    12-13




  • Unstructured Data II: Information Extraction Methods and Knowledge Representation



    Getting to Start with Examples of Intelligent Systems (AI) with Big Data, Big Data Processing and Big Data Analytics:



    Knowledge Representation for AI

    Example of Knowledgebase Based AI Application:

    LectureNotes: Overview of Knowledgebase Based Question Answering System (IBM Watson)






    Knowledge Representation for AI


    Big Data: Unstructured Text Data Processing for Information Extraction

  • Completely Unstructured Text Documents -- Extracted Webpage Content, Social Media Messages, Electronic Books (in pdf)

  • Introduction to Text Processing for Information Extraction and Knowledge Representation of Text Documents:



    Important Big Data Processing Techniques: Information Extraction(IE) Methods:






    Information Extraction Methodology for Unstructured Text Documents


  • How to Represent Knowledge in Texts

  • Unsupervised Information Extraction Methods

  • Open IE (Information Extraction) Parser

  • OpenIE Example from Stanford NLP Research site

    Stanford NLP Open IE (Information Extraction) for Triple Extraction
    Core NLP Parsers from Stanford NLP Research

    Demo site: Core NLP online Test Run Site (Stanford)






    Triple Extraction Method Using Stanford OpenIE (Information Extraction)

    Sample Application of Knowledgebase based Medical QA :

    Document Data Source: Medical Information Site:
    Medline Plus Site for Medical Encyclopedia

    Building a Knowledgebase Based a Medical Question Answering(QA) System Using OpenIE Method

    Example of Big data Project with Open IE Method: Building a Medical Question Answering System

    Lecture:
    How to Build a Knowledgebase from Medline Plus Site of Medical Encyclopedia using OpenIE



    Lecture Notes on Medical QA: Entity-Relation Extraction Methods in IBM Watson Medical KG (From the Paper:IBM Watson)







  • How to Represent Knowledge in Wiki Webpages

  • Lectures:

  • Semantic Web:RDF (Resource Definition Framework) for Representation of Knowledge in Wiki Page Contents

  • RDF: Graph to Represent Knowledge in Web Page Contents in TRIPLES (Subject, Predicate, Object):

    Semantic Web: RDF (Resource Definition Framework) to to Represent Knowledge in Web Page Contents


    Semantic Web RDF Data Model - Semi-Structured Data Model

  • RDF is a Graph Data Model
  • RDF Data are Directed, Labeled Graphs
  • A Single Edge Connecting a Node to another Node in an RDF Graph is a 3-tuple that is called either a Statement or Triple
  • Triples are organized into named graphs, forming 4-tuples, or quads
  • RDF resources (nodes), predicates (edges), and named graphs are labeled by URIs
  • Semantic Web technologies including Web Ontology Language (OWL) and SPARQL for RDF Graph Database Query Language


  • Lecture Notes: Intro to RDF for Semantic Web ***************

    Lecture Note on RDF Scheme (RDFS) for Semantic Web


    W3C Site for Registered Organization Vocabulary (rov) in RDF for Wikidata


    Tutorial to Understand Basic Concept of RDF for Semantic Web

    Example Visualization of RDF Graph Database for RDF Concept Vocabulary for Wiki Data

    Tutorial Lecture Semantic Web in RDF to Represent Web Contents Youtube Video





    Semi-structured Data Query Language

    XPATH:

    Lecture Notes On XPATH

    XQUERY:
    Lecture Notes On Semi-Structured Databse Query Languages: XQuery



    SparkQL: RDF Query Language for Retrieval of Knowledges in a RDF Databases

    Lecture on Intro to SparkQL: RDF Query Language for Semantic Web ***************


    Real World RDF Application on Wiki Site Web Data

    Tutorial on SparkQL over RDF for Wikidata






    Example of AI Application Project: Medical QA System Using RDF Graph Database and SparkQL Query Language:

    Lecture Note on AI Application: Medical QA System Using RDF and Semantic Web Technologies



    Related Recent AI Research Paper:

    Universal QA Platform for Knowledge Graph (ACM SIGMOD 2023)





    Official W3C RDF Related Sites:

    Official Resource Definition Framework (RDF) XML RDF Guide in W3C

    W3C RDF Specification

    Official SparkQL Guide Site in W3C

    Official SparkQL in N3 (instead of XML) Guide Site in W3C














  • Knowledge Graph (KG)


  • How to Represent Knowledge in Text Documents for ML/AI Algorithm to Learn to Answer?

    Information Extraction Methodology for Unstructured Text Documents


    Unsupervised Information Extraction Methods

  • Open IE (Information Extraction) Parser

  • OpenIE Example from Stanford NLP Research site

    Stanford NLP Open IE (Information Extraction) for Triple Extraction
    Core NLP Parsers from Stanford NLP Research

    Demo site: Core NLP online Test Run Site (Stanford)






    Triple Extraction Method Using Stanford OpenIE (Information Extraction)

    Sample Application of Knowledgebase based Medical QA :

    Document Data Source: Medical Information Site:
    Medline Plus Site for Medical Encyclopedia

    Building a Knowledgebase Based a Medical Question Answering(QA) System Using OpenIE Method

    Example of Big data Project with Open IE Method: Building a Medical Question Answering System

    Lecture:
    How to Build a Knowledgebase from Medline Plus Site of Medical Encyclopedia using OpenIE






    Foundation Papers on KG to Build MedQA:


    Paper IBM Watson for Medical QA: Entity-Relation Extraction Methods in IBM Watson Medical KG 2015

    Paper: Knowledge Graph for Medical QA: Entity-Relation Extraction for Medical KG (IAAA 2017)

    Paer Presentation: IBM Watson for Medical QA: Entity-Relation Extraction Methods from the Papers IBM Watson Medical KG

    Paper Presentation: on Medical QA: Entity-Relation Extraction (from the Papers IBM Watson Medical KG and IAAA2017)







  • Entity Extraction
  • Stanza: Stanford Unsupervised Stanford Medical NER (Named Entity Recognition)

    Performance of Stanza: Unsupervised Medical NER (Named Entity Recognition)

    How to Use Stanza CRAFT Biomedical NER Types (stored as property for each entity in triple) to Process a Question better:

    Stanza CRAFT: Unsupervised Stanford Biomedical NER (Named Entity Recognition)


    Papers:
    Stanza CRAFT Paper (2021): Biomedical Named Entity Recognition

    Recent Named Entity Recognition 2024



  • Relationship Extraction
  • Papers:
    REBEL: Relation Extraction By End-to-end Language generation



    Good Tutorials for Information Extraction Methods to Build KG

    Information Extraction (IE) Methods in Triples for Knowledge Graph(Using SpaCy)










    Knowledge Graph (KG): How to Build Knowledgebase using Graph Database

  • Graph Database Model

  • Property Graph Data Model - Semi-Structured Data Model

  • Property Graph is a Semi-Structured Graph Data Model
  • Property Graph Data are Directed, Labeled Graphs
  • A Single Edge in a Graph is a 3-tuple that is called triple
  • Triples are organized into Named (Labeled) Graphs, forming [Subject Entity Node, Relationship (Edges), Object Entity Node] with Property
  • Each Node and Edge Can Have Multiple Properties


  • Property Graph Database Server: Neo4j

  • Neo4J Download:
    Neo4j Graph Database Download


    Intro to Graph Database Concept

    Neo4j: Graph Database Model



  • Cypher Graph Query Language Based on Property Graph Model

  • Lecture: Property Model Based Graph Query Language: Intro to Cypher

    Intro to Cypher Graph Query Language

    Cypher Graph Query Language Comparison to SQL
    Cypher Query Pattern


    Tutorial: Cypher: Graph Modelling with Label

    Lecture: CypherQL on Retrieving by Relationship in Triples


    Playing with Cypher:
    Tutorial: Starting with Neo4j Graph Database with Cypher Query Language


    Comparison from SQL to Cypher

    How to Import from Realtional Table to Graph

    How to Import from JSON (NOSQL) to Neo4j Graph



    Related Research Paper:
    Cypher Graph Query Languane on Property Graph Model (ACM SIGMOD 2018)





    Research on Knowledge Representation for AI

    Building a Medical Knowledgebase in Medical Knowledge Graph (KG) for Medical QA:

    Lectures:

    Lecture Notes on Knowledge Graph: Unsupervised Entity Relation Extraction in KG Text Processing (From the IAAA 2017 AI Paper)

    Lecture Notes on Knowledge Graph: Unsupervised Entity Relation Extraction in KG Text Processing (From the KDD 2019 AI Paper_

    Lecture Notes on Knowledge Representation in Medical Knowledge Graph (From the Big Data Lab Research Paper)

    Lecture Notes on Building a Medical Question Answering (QA) System with Medical KG (Big Data Lab Research)




    Related Papers: (Presentation for Extra Credit at the Final Project Presentation)

    Unsupervised Knowledge Graph with Role of Conditions for Text Processing KDD 2019











    Current Research on QA System with Generative AI: Text-to-SQL

    Amazon Generative AI: Text-to-SQL Answering Sysetm

    Paper: Text-2-SQL Survey VLDB Journal 2023



    Current Research on QA System with Generative AI: Text-to-Cypher

    Paper: Generative AI: Text-to-SPARQL


















    13-14




  • NLP with Machine Learning: Large Language Models (LLMs)





    Natural Language Processing (NLP) with Machine Learning



    Review:

    Problems (Limitations) with Lexicon Based TF-IDF for Text Analysis:

    Phrase (N-Gram Word) Identification
    Negation Handling
    Can NOT Identify Relationships among Terms - Synonyms (Similar Meaning) or Antonyms (Opposite Meaning) of Terms
    Can NOT Identify Different Meanings of a Same Term by the Different Context of a Senetence - Polysemy Problem !
    Can NOT Identify New Terms, Common Slangs or Changining Relationship Between New Terms (ex: Data Analytics) and Old Terms (ex:Data Mining)
    Can NOT Identify Sarcastic Contexts
    Can NOT Identufy the Order of Words in a Sentence !!!!


    Problems of Learning a Sequence of Tokens in a Sentence/Phrase

    Old Method to Learn Word Sequence: Positioning Index for Phrase Query (Stanford)













    Semi-Supervised Learning for Text Analysis:

    Context Aware NLP Methods Using Machine Learning with Semi-Supervised Learning: Word2Vec (Google)

    Word2Vec:

    Lectures:

    Tutorials for Word To Vector: Skipgram Model for Training Word Sequences in a Phrase/Sentence

    Theory to Understand How Word2Vec works with Skipgram Training Model (Advanced)



    To get an Executable Binary of word2Vec Model Implementation and Training Data sets:

    Google AI Implementation sites of Word2Vec 2013:
    Google Original Sites for Word2Vec

    Original Word2Vec Papers by Google:

    Google Original Paper for Word2Vec
    Google Paper for Optimization for Word2Vec


    Extended Word2Vec: Glove from Stanford:
    Sites for Glove Word Vector by Stanford NLP Team

    Tutorial Sites for Word2Vec Implementation:

    Tutorials for Word To Vector with Google Tensorflow
    Tutorials for Word Embeddingd with Google Tensorflow


    Other Related Sites

    Sites for Word To Vector for npmjs
    CNTK: Stanford NLP Sentiment Analysis



    Large Language Model (LLMs) - NLP Deep Learning Models

    Google BERT (Bi-directional Transfomer)

    Lecture on the First LLM: Google AI BERT (Bi-directional Encorder Representations from Transformers)

    What is Google AI BERT Transformer?

    BERT: What this means to your Web Search?

    SpaCy Embeddings for BERT Transformer





    Early Basic Papers to Understand Passage Ranking Methods for Document Retrieval and Learning

    Research Paper: Passage Construction and Ranking Methods for Question Answering System (ACM 2018)
    Research Paper: Passage Ranking Approach for Document Learning
    Research Paper: Passage Reranking Using BERT from Google (2019)











  • RAG (Retrival-Augumented Generation): Document Vector Databases (by Facebook MetaAI)

  • How to Represent Knowledge in Text Documents for ML/AI Algorithm to Learn to Answer?

    Retrieval-Augmented Generation (RAG):

    Lectures:

    Lecture Notes on FAISS and DPR: Intro to RAG for Building a Document Vector Knowledge Database for Retrieval


    Original RAG Papers by Facebook Meta AI:

    Research Paper (Facebook Meta AI): Retrieval-Augmented Generation (RAG) for Knowledge-Intensive NLP Tasks

    Research Paper (Facebook Meta AI): Dense Passage Retrieval (DRP) for RAG

    Research Paper (Facebook Meta AI): FAISS for RAG


    Overview of Current RAG Based Approaches and Research Trends with Knowledge Graph Based AI (QA Systems)


    Tutorial: LangChain framework to build QA with LLMs


    More on RAG:

    RAG Framework to build QA with LLMs
    RAG Papers









    Current Research on RAG Based QA System with Generative AI: Text-to-SQL

    Amazon Generative AI: Text-to-SQL Answering Sysetm





    13-14



  • Lectures on Massively Parallel Distributed Big Data Processing Systems


    Big Data Processing Systems and Parallel Distributed Programming Paradigm


    Review on Performance in File System of RDBMS:

    Class Note_8: Chapter 17 Disk Storage System, File Structure, Hash Index
    Class Note_9: Chapter 18 Index Structure for File




    Massively Parallel Distributed Big Data Processing: MapReduce with Hadoop Distributed File System (HDFS)

    Motivations:

  • Building Data Pipelining for Information Extraction Using Parallel Semistructured (NoSQL) Data Processing Systems
  • Examples of Industry Big Data Processing Systems: LinkedIN (and Example of Final Big Data Project)

    Examples of Data Pipeling for Query Execution Steps in Big Data Processing Systems

    Examples of Industry Big Data Processing Systems at Uber


    Overview of Parallel Data Processing Systems for Big Data Processing:

    Introduction to Big Data, Big Data Processing, and Big Data Processing Systems On Cloud
    Overview of Big Data Processing Systems with NO SQL Systems


    General Architecture of Massively Paralle Processing (MPP) Systems as a Platform as a Service (P-a-a-S) in Cloud


    MapReduce by Google:

    Lecture Notes On Map Reduce (Google)

    Original Paper: Map Reduce by Google






    Hadoop Distributed File System (HDFS) - Implementation of Google MapReduce by Apathe and Yahoo:

    Lecture Notes Part I on Introduction to Hadoop Distributed File System (HDFS)

    Lecture Notes Part II on MapReduce Executions on Hadoop Distributed File System (HDFS)


    External Sort and External Hash Partition on Single Node (From CIS611 Notes)



    Hadoop Job Class

    Hadoop MapReduce Tutorials:

    Hadoop Map Reduce Tutorial (New)
    Map Reduce Tutorial on Hadoop

    Hadoop Interface Input Format - Input Split: getSplits()
    Hadoop Interface for Writable
    Hadoop Interface for WritableComparable
    Hadoop Interface for OutputCollector
    Hadoop Interface for Map
    Hadoop Interface for Comparator to control Grouping
    Hadoop Interface for Combiner to Local Aggregation on the intermediate Output
    Hadoop Interface for Reducer

    Hadoop Parameters: io.sort.mb


    Hadoop Implementation Set Up:
    Hadoop Single Node Set Up
    Hadoop Fully Distributed Mode (Cluster) Set Up


    Code Examples of MapReduce on Hadoop:
    Simple MR Hadoop Lab Guides
    MR sample Codes - Average Temp By Zip Code
    MR Sample Codes - Sort By Station ID
    station Data Set







    15


  • Distributed Parallel Big Database Processing Systems on HDFS:


  • Quick Review: What are Data Warehouse and OLAP (Online Analytical Processing) ?

    Only For Those Who Have Not Taken CIS 611 Enterprise Database Systems with DW and OLAP. (DW and OLAP won't be covered in this course)
    Class Note_24: Lecture Notes On Overview of Data Warehouse with OLAP
    Class Note_24_1: Paper on Data Cube by Jim Gray (Microsoft), et al


    Motivation:

    LinkedIn: Industry Research Paper: Building a Real time LinkedIn System on Social Media Platform with Data Warehouse Approach
    Avatara: OLAP for Webscale Analytics Products (LinkedIn)

    Example Projects for Building Big Data Processing System for Data Analytics with Social Media Big Data

    Example of Big Data Project
    Project: Data Mining over LinkedIn Data





    Lectures:

  • Massively Parallel Processing (MPP) Systems: NOSQL Systems

    Lecture Notes On Introduction to NOSQL Systems



  • Hive:

    Class Note_19: Lecture Notes On Hive (Facebook)

    Paper: Hive by Facebook

    Hive Documentations:
    Apathe Hive Site
    Apathe Hive Documentation
    Apathe Hive Tutorial

    Important Hive Tutorials:
    Apathe Hive Getting Started
    Hive Lanuage Manual: DDL, DML

    Examples of Hive DDL with Partition By:

    Example of Creating a Hive Table with Partitioning
    Required Setting for Partition and Step by Step Example for Creating Hive Partitioning Tables with Partition By
    Note that For Craetion of Partition Tables, for partition, you have to set this property set hive.exec.dynamic.partition.mode=nonstrict

    Hive Data Types
    Hive Complex Data Types
    Hive Complex Data Types

    Code Examples of HIVE DDL with Columns of Complex Data Types
    Examples for Hive Complex Data Types: Array, Map as a Collection of key:value pairs

    Transform Clause to Provide User Customized Map and Reduce Script

    Hive SerDe:
    Guide-Hive SerDe (Serializer and Deserializer)
    Hive Row Format with SerDe (Serializer and Deserializer)
    Hive DDL-Row Format, Storage Format,and SerDe
    Hive Built in SerDe

    Code Example for Hive CSV SerDe
    Code Example for Hive JSON SerDe for Twitter data
    Code Example for Hive JSON SerDe for Twitter data (in doc)

    UDTF:
    Hive UDTF Tutorial
    Hive UDTF Example
    Hive File Format

    Sample Project Using Hive:
    Sample Project: Amazon Review Data Processing and Analytics Using Hive and JSON SerDe
    Sample Project: LinkeIn Data Processing and Analytics Using Hive
    Sample Project: Twitter Data Processing and Analytics Using Hive and MongoDB




  • Apathe Pig Latin:

    Class Note_20: Lecture Notes On Pig Latin (By Yahoo)
    Class Note_20_1: Lecture Notes On Pig Latin (By Yahoo)

    Paper: Pig Latin By Yahoo

    Class Note_20_2: Lecture Notes On MapReduce Join Algorithms
    Class Note_20_3: Lecture Notes On Advanced Partitioning Techniques

    Apathe Pig Latin Site to Start

    Sample Project Using PIG Latin:
    Sample Project: Nasa Network Server Log Data Processing and Analytics Using PIG Latin and R



  • Apathe Parquet: Columnar Store System on HDFS (CIS611 Covers a Fundamental Columnar Concept and Columnar Systems)

    Apache Parquet is a columnar storage format available to any project in the Hadoop ecosystem,
    regardless of the choice of data processing framework, data model or programming language.

    Apathe Parquet Site




  • MongoDB: Advanced on HDFS

    Class Note_24: Lecture Notes On MongoDB

    MongoDB Expressions for Conditions:
    Mongo DB Query Selector and Projection Operators
    Mongo DB Condition Using $expr With Conditional Statements
    Mongo DB Logical Operators
    Mongo DB Condition with Logical Operator $and

    MongoDB Aggregate Pipelining:-- Advanced
    MongoDB: Aggregation Pipelining with MapReduce
    Complete List of MongoDB Aggregate Operators for Pipelining
    MongoDB: To start for Aggregation Pipelining
    MongoDB Add Field with Aggregate operation Pipelining
    MongoDB Aggregate operation Pipelining for Bucket


    MongoDB lookup Operator for Join with Aggregation Piplelining :

    MongoDB Aggregation Pipeline Operator List
    MongoDB $lookup operator for Join: Join returns null matching as well (Left Outer Join Semantics)
    MongoDB $lookup Join operator with $unwind for Array Valued Field as Join column
    MongoDB $lookup operator with $let for Multiple Fields for Join Columns
    MongoDB let operator

    MongoDB CRUD -- Advanced:
    MongoDB UpdateMany Operation with Aggregation Pipelining and $set as $addfields, $unset Operators
    Mongo DB: Update with Upsert = true
    Mongo DB: Bulk Write


    Mongo DB Index, View, Distributed Feature- Sharded Collection:
    MongoDB Index
    MongoDB Compound Index
    MongoDB Text Index Features
    Mongo DB: View for Read Only
    Mongo DB: BSON Document Format
    Mongo DB Distributed Cluster: Sharded Collection

    MongoDB Concurrency Control:
    Mongo DB: Update with Upsert = true
    MongoDB Update with upsert and Controlling Concurrency Probelms

    MongoDB Other Operators:
    Mongo DB: text serach Operator
    MongoDB System variable
    Compelete List of Mongo DB Operators

    How to Discover MongoDB Scheme:
    How to Discover MongoDB Schema Information
    Mongo DB schema-analyzer
    Mongo DB schema-analyzer Tool: Variety

    Sample Project Using MongoDB:
    Sample Project: Twitter Data Processing and Analytics Using MongoDB

    Mongoose for connecting to MongoDB with Node JS Express:
    Mongoose To Set Up
    Mongoose Guide
    Node API for MongoDB:
    NodeJS: Integrating to MongoDB

    MongoDB Sites to Set Up:
    Class Note_22_3: MongoDB Site
    Mongo DB Resources
    MongoDB Download
    Mongo DB Getting started
    Mongo DB Documentation
    Mongo DB Glossary: _id




  • Google's Big Table:
    Class Note_21: Lecture Notes On Big Table (By Google)
    Paper: BigTable by Google



  • FaceBook's HBase:

    Class Lectures:
    Class Note_23: Lecture Notes On HBase (FaceBook)
    Class Note_23_1: Lecture Notes More On HBase Architecture
    Class Note_23_2: Apathe HBase Hive Integration (by HortonWorks)
    Apathe HBase Hive Integration (in MapR)

    Quick Start Guides:
    Class Note_23_3: Quick Guide On HBase
    HBase Basic Tutorial to start
    HBase with Hive Set Up guide by Paul Webster
    HBase Tutorial with HBase Join with Hive by Paul Webster

    Apathe HBase Sites to Dive:
    Apathe HBase Site To Start
    Apathe HBase serDe
    Apathe HBase Hive Integration


    Major Prebuilt Platforms for Hadoop with NoSQL Stack:
    Hornworks:
    Hornworks Big Data Platform for Prebuilt Hadoop Stack with HBase Hive Integration
    Cloudera:
    Cloudera Big Data Platform for Prebuilt with Spark Hive Integration
    Cloudera Big Data Platform for Prebuilt Data Scientist Work Bench
    Cloudera Set Up for Spark Tutorial From Danielle Aring's Report

    Comparison: Hornworks vs Clodera vs MapR for Prebuilt Hadoop Stack Platform with HBase Hive Integration


    More on HBase:
    Apathe HBase Architecture
    More in Depth of Apathe HBase Architecture
    Apathe HBase Architecture

    Papers on HBase:
    Data Warehousing and Analytics Infrastructure at Facebook
    Big Data Processing System at FaceBook (From SIGMOD 2013)
    Paper on HBase
    Paper on HBase



  • Cassandra:
    Class Note_25: Lecture Notes On Cassandra
    Apache Cassandra Site to start




  • Real Time Big Data Stream Processing System:

  • Spark: In Memory Processing for Streaming

  • Class Note_26_1: Lecture Notes On Spark: Spark Streaming/SparkSQL/SparkR


    Spark Platform Components:

    Spark RDD and Operations
    Spark MLLib Guide
    Querying Hive Table in SparkSQL

    Quick Tutorial to Start:

    Berkerley Spark Tutorial (in Scala or Java)
    Spark Tutorial in Python Examples
    Spark Lamda Function in Python Example
    Spark Tutorial in Python Examples more


    PySpark Platform :

    PySpark for Beginners
    PySpark Tutorial for Beginners
    PySpark Tutorial
    PySpark Tutorial for Data Science



    Spark Tutorial with Text Analysis Project Examples:
    Spark Tutorial I for Set Up on Cloudera VM and Basic Text Analysis (From Danielle Aring's Report)
    Spark Tutorial II for Product Review Text Analysis (From Danielle Aring's Report)


    Prebuilt Data Analytic Platform with Cloudera:

    Cloudera Big Data Platform for Prebuilt with Spark Hive Integration
    Cloudera Big Data Platform for Prebuilt Data Scientist Work Bench
    Cloudera Set Up for Spark Tutorial From Danielle Aring's Report

    Sample Project Using Spark and MapReduce:
    Tutorial:Real Time Twitter Analysis with Spark and Kafka
    Analysis of State Union Address with We vs Me index
    Analysis of Amazon Review Data with Spark

    Useful Spark Sites:
    Apache Spark Site to start
    Apache Spark Documentation
    Spark Programming Guide
    Spark with Cloudera Site to start
    Launching Spark on Amazon EC2 Cluster

    Berkerley Spark Tutorial (in Scala or Java)
    Spark RDD
    How to Get Twitter Data

    Spark Twitter Collection in Scala
    Spark Pipeline
    Spark Pipelining Example
    Spark RDD Partitioning
    Spark Reduce and Fold Example

    Papers:
    Apache Spark for Stream DM
    Stream Data Analytics with SparkSQL
    Apache SparkSQL
    Apache SparkR
    Shark
    Shark SQL


    Lecture Notes_27 On Storm: Another Realtime Data Stream Messaging System

    Apache ZooKeeper







    kafka: Asynchronous Distributed Big Data Stream Messaging System on Cluster

    Lectures:

    Introduction of Apathe Kafka
    What is kafka?

    Apache kafka Documentation

    Apache kafka Streams Core Concepts

    Architecture of Apache kafka Streams


    Tutorials:

    Apache kafka Streaming Data with Example Codes

    Apache kafka Streaming Tutorial with Codes

    Apache kafka Developers Guide


    Intro of big data messaging with kafka
    kafka Set up Instruction Note that you need Java and Zoo Keeper for Kafka Set Up

    Sample Project Using Kafka:
    Real Time Twitter Analysis with Spark and Kafka





  • Parallel Big Data Processing Systems: New SQL Systems

  • VoltDB:
    Class Note_28: Lecture Notes On New SQL
    Class Note_29: Lecture Notes On VoltDB (New SQL)
    VoltDB Site to start











  • 15


  • Transaction and Concurrency Control

    Transaction in Relational Database System (RDBMS):
    Lecture Notes on Transaction and Concurrency Control in RDBMS:

    Class Notes on Overview of Relational DBMS with Transaction and Concurrency Control

    Review: Examples on Transaction with Multiple Concurrent Users


    Class Note_4: Lecture Notes On Transaction
    Class Note_5: Lecture Notes On Concurrency Control









  • 15



  • Big Data Processing Systems on Cloud Computing

    Class Note_30: Lecture Notes On Cloud Computing
    Mobile, Multimedia and Cloud Computing


    Amazon Cloud Service:

    Amazon Elastic Cloud Computing (EC2) for Web service

    You can use a Container: Docker (Instead of VM) to Configure a Distributed Cluster

    How to Create Amazon Virtual Hosting (EC2) with a Web Server and Amazon Database Server (RDS)
    Amazon DynamoDB Database Service
    Amazon RDS (Relational Database Service)

    You can create your Big Data Processing Infrastructure on Amazon Cloud
    Amazon Cloud Account for Students


    Microsoft Cloud Service Azure

    Trial account for Microsoft AZURE Cloud
    AZURE Cloud
    Tutorial for MS AZURE Cloud
    Sample Project on Microsoft AZURE Cloud
    How to Create/Retrieve a Table in Microsoft AZURE Cloud
    How to Create/Retrieve BLOB data in Microsoft AZURE Cloud
    How to Create a SQL Database Server in Microsoft AZURE Cloud



  • 6-7


  • Big Data Platform in Web Applications

    Web Server Programming with Node JS and Mongo DB (MEAN Stack)


    Serverside JavaScript Framework: NODE JS

    Class Note_18: Controller Server Communication
    Class Note_19: Introduction to Web Server
    Class Note_20: Introduction to Node JS
    Class Note_21: Express

    Good Node JS/Express Tutorial Site
    Express Basic Routing
    Angular route Parameters



    NodeJS Set Up Guide

    NodeJS Site
    NodeJS Download
    NodeJS Package Manager Download
    NodeJS Installation Tutorial on Window
    NodeJS Installation Tutorial on MacOS
    NodeJS Installation Tutorial on Ubuntu/Linux

    Another NodeJS new source site for NodeJS Set Up
    NodeJS Download



    Node JS Framework: Express
    Node JS Framework: Express
    Express Set Up: Getting Started

    Express Coding Guide:
    Express Example: How to Create a Simple App
    Express Tool to Generate an Express App
    Express Basic Routing
    Angular route Parameters
    Express Guide: Routing to Write
    Node JS HTTP Methods to Write

    Basic Express Built In Frameworks/Middleware:
    Node JS Documentation: Global variables
    Node JS Documentation: _dirname, _filename
    Express Guide: API Methods: res.json()
    Express Guide: Writing a Server Side Code:Middleware
    Express Guide: Integrating a Database
    Express Guide: Static File Server
    Node JS: server.address()
    How to Know your server address on Window
    Node JS: multer for file upload
    Node JS: bodyparser.urlencoded with extended false/true
    Node JS: bodyparser.urlencoded with qs
    Node JS: bodyparser.urlencoded

    Node JS: request object
    Node JS: Node JS: response object
    Node JS: setting response object
    Node JS: http API



    Other Resourses:
    Node JS Documentation
    Framework Built on Node JS Express





    NODEJS SAMPLE APPLICATION CODES

    Web Server Code Examples with NodeJS and Express :

    Good Node JS Express Tutorial Site

    Tutorial for Express Basic Routing

    First Node App






    Integrating a Database Server with Node JS: MySQL/MongoDB/PostgreSQL/SQL Server:

    Class Note_22: Database Server

    Node API for MySQL:
    NodeJS: Integrating to MySQL
    Node-MySQL


    NodeJS works with Most of Databases:
    Example Codes for Connectivity to Each Database in NodeJS

    Database Connectivity in NodeJS

    Node Package Manager (NPM) for Database Connectivity

    Microsoft SQL Server: Database Connectivity in NodeJS
    MySql: Database Connectivity in NodeJS
    Postgre: Database Connectivity in NodeJS



    Application Code Examples with NodeJS with Database Server MongDB or MySql :


    Simple Node Application Codes with MongoDB

    Node Application Codes in MVC with MongoDB

    Node Application Codes in MVC with MySql





    Set up Guide for MEAN Stack: Also See Project 3 Section For Node JS Set up Guide
    Node JS with Mongo DB Setup Guide
    Node JS with Mongo DB Setup Guide
    Angular JS with Node JS with Setup Guide

    Node JS API for MongoDB:
    NodeJS: Integrating to MongoDB
    Mongoose: Node JS API for MongoDB: For connecting/querying to MongoDB:
    Mongoose To Set Up
    Mongoose Guide


    Application Examples Built with Node JS
    Sample Web Application Using Node JS with Mongo DB
    Sample Web Application with Angular JS and MS SQL Server



  • 16

    Project Presentation
    Project Presentation Scedule

    ==> Completion of Homeworks/Labs is required for obtaining a passing grade.

  • This is a tentative scale and
    it could be changed

    Letter
    Grade

    Quality Points

     


      A

    > 93%  

    A: Outstanding (student's performance is genuinely excellent)

      A-

    90% - 93%

     

      B+

    87% - 90%   

     

      B

    82% - 87%

    B: Very Good (student's performance is clearly commendable but not necessarily outstanding)

     

      B-

    80% - 82%

     

     

      C

    75% - 80%

    C: Good (student's performance meets every course requirement and is acceptable; not distinguished)
        D 65%-75% D: Below Average (student's performance fails to meet course objectives and standards)

     

      F

    <65%

    F: Failure (student's performance is unacceptable)

    ADA Adherence. If you need course adaptations or accommodations because of a disability, if you have emergency medical information to share with me, or if you need special arrangements in case the building must be evacuated, please make an appointment with me as soon as possible. My office location and hours are listed on top of this syllabus. If you need further information, please contact the ACCESS office, phone number 687-5106.

     


    Programming standards

    • Every program must include your name, CSU ID number, Class, Section Number, Hours, the words 'Homework # ...', and a short description of the assignment. For example:
       ' Name: Mark Zuckerberg  
       ' ID: 1234567            
       ' Homework #1            
       ' Description: Computing the average life of a light bulb
    • Every variable should have a meaningful name (this includes function/procedure/subprogram names).
    • Every portion of the program should be as cohesive (single purposed) as possible. This leads to a large number of small functions.
    • Every function (including the main function) should be preceded by a comment indicating its arguments and a description of the transformation it performs.
    • Non-obvious code within a function should be explained.
    • Code should not be over commented.