CIS 612 Big Data and Parallel Distributed Data Processing Systems (3-0-3) |
||
| Course Content |
|
|
|
10. Jan 14, 2026: The CIS612 Class Time Will Be 4:15PM - 5:30PM During This Spring Semester! Midterm Info: //a href="https://eecs.csuohio.edu/~sschung/cis612/CIS612Spring26_MidtermLinksOnly.html"> Information on Midterm and Midterm Lecture Links Only (Scroll Down to the Lecture Note Sections) The Topics to Focus On for the Midterm: IMPORTANT WARNING !! If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterm and the Final !! Characteristics and Architecture of Big Data Processing Systems/Applications, Cloud Computing Systems Universal Data Exchange Formats in Object Exchange Model(OEM): Three Common Big Data formats in Semi-Structured Model: Data Model, Syntax of Each Encoding Format, the Related Data Processing Techniques: DOM, XPath Characteristics of Semi-Structured Data Model, Main Differences between Semi-structed Data Model and Relational Data Model for Object Relation Mapping (ORM), Conversion among Relation(CSV), XML, and JSON Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining Comparison Between Semi-structured Database Server and Relational Database Server for Database Management Strategies NOTE that ONE Page Note is NOT allowed to the Midterm !! Exam Formats: 6-7 Main Questions with 2-3 sub questions for a Small/Short Answer Types for Problem Solving Final Exam Information: Monday May 4th 4:00PM-6:00PM In Each Lecture Note for the Final, Only the Slides (Subjects) Covered in Class Will Be in the Final See the Lecture Notes Links Only for the Final here (All the rest of links are disabled) //a href="CIS612_FinalLinksOnly.html"> the Lecture Notes Links Only for the Final One Page Note (Each Side) Is Allowed to The Final. It Shoule Be Hand-Written. Printed or Copy and Paste of Entire Letcure Notes Are NOT Allowed. Topics to Focus on for the Final Unstructured Data Processing, Inverted Index, How to Build Inverted Index, TF-IDF Based Ranking Algorithm with Cosine Similarity Text Cleaning Preprocessing, Basic Natural Language Processing (NLP) Methods in Data Pipeling, POS Tagging Problems of TF-IDF Based Ranking for Text Analysis Context Aware NLP Methods - POS Tagging, NER Tagging Design of Intelligent System with Data Pipelining of Big data processing Problems of TF-IDF Based Ranking for Text Analysis Context Aware NLP Methods - POS Tagging, NER Tagging, OpenIE Design of Intelligent System with Data Pipelining of Big data processing Information Extraction Methods, Design of an Answering System, How to Build a KnowledgeBase from a Large Collection of Unstructured Texts for Question Answering System Semi-Structured Query Model - XQuery Semi-Supervised Learning - word2vec for Question-Answering System XQuery Using XPath Intro to RDF Concepts Graph Database Neo4j Cypher Query Basics Parallel Map Reduce Programming, Parallel Distributed Data Processing Execution Algorithm in MapReduce (the Execution Phases (Steps) of Map Reduce on HDFS), Data Excution Algorithm Steps in Map phase and Reduce phase, Architecture of HDFS -- Architecture of Parallel Distributed File System in General Architecture of Parallel Distributed Platform (Amazon EC2 Cloud Platform in General) Hive DDL to Build DW like Hierarchical Directory Structures Pig Latin Data Pipelining Execution Steps for General Analytic Queries One Page Hand Written Note (Each Side) is Allowed to the Final ! Printed Copy and Pasted Lecture Notes Are NOT Allowed as One Page Note 9. Jan 12, 2026: Check the Final exam schedule and the Semester Schedule below //a href="https://www.csuohio.edu/registrar/academic-calendar"> University's Official Academic Calendar for the Semester Schedules to Add, Drop, Withdraw, and the Final Exam schedules 8. Jan 12, 2026: It is Required to Attend Every Class ! Your Class Attendance Will be Checked in Each Class. There Will be Random Quizzes to Check the Attendance ! Only 5 - 10 Mins Will Be Allowed for Each Quiz. Those Who Come in Class Late After 5 Mins, They Will NOT Be Given a Quiz ! 7. Jan 12, 2026: The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class Blackboard ! Wait For the Lab Submission Links To Be Created on Blackboard with the Deadline to Submit Each Lab 6. Jan 12, 2026: Only the registered students can access the course blackboard. If you have a problem with your blackboard access, please contact the registrar and CSU e-learning Tech Support to resolve the issue ! Faculty do not control your registration in the CSU Campus Systems and the course blackboard access. //a href="https://www.csuohio.edu/center-for-elearning/technical-support"> e-learning CSU Tech Support 5. Jan 12, 2026: TA Information: TA: Email: k.chau@vikes.csuohio.edu Office Hours: Mon, Wed 1:00-3:00 PM (Officially); Tue, Thus 2:00-4:00 PM (if needed) Send email to TA ahead to set up a time slot to let him know that you are coming Location: Big Data Lab: FH 305 or a ZOOM Meeting ZOOM Meeting ID: 758 582 3674 Passcode: //a href=""> If you have questions in the Labs or Grading your Labs, Send an email to TA to See During the TA's Office Hours or Schedule a meeting in the TA's Zoom meeting Dr. Chung's Office Hours: Tues and Thurs 1:30PM - 3:30PM Send Me Email Ahead to Set Up an In-Person or Zoom Meeting. Email: s.chung@csuohio.edu Zoom Meeting Info: Meeting ID: 859 3867 6332 4. Jan 12, 2026: Lab Submission: The Output of each lab is your report in Doc file that shows your screen captures of the Executions with Each Output (each intermediate output as well as final outputs) Generated. your Report in Doc file also should explain all the platform set up, the execution steps, and copy of each source code files ) Each of your screen capture must show your results returned in YOUR SYSTEM to prove that your lab is done correctly by YOU !! Submit your Zip file that includes the followings on Blackboard for a timestamp and as a proof 1) Your Lab Report in .doc file that explains all the platform set up procedures, the execution steps, each intermediate output, final outputs, and a copy of each source code files, 2) All your Source files, and output files 3. Jan 12, 2026: Each Class Attendance is Required for CIS612 ! There Will Be Random Quizzes and Sign Up Sheet to Check Each Attendance of You ! 2. Jan 12, 2026: You can use any SQL Server for Your Database such as MySql or MS SQL Server. See Set up Guide for MS SQL Server or MySQL in the Lab Section. 1. Jan 12, 2026: The class webpage for this semester was announced on the class blackboard: //a href="https://eecs.csuohio.edu/~sschung/cis612/CIS612Spring26.html"> CIS612 Big Data Class Webpage Check the Last Day to Add and Drop here ! //a href="https://www.csuohio.edu/registrar/academic-calendar"> University's Official Academic Calendar for the Semester and the Final Exam schedules CIS612Final Exam: Monday May 4 4:00p-6:00p Midterm Info: //a href="https://eecs.csuohio.edu/~sschung/cis612/CIS612Spring26_MidtermLinksOnly.html"> Information on Midterm and Midterm Lecture Links Only (Scroll Down to the Lecture Note Sections) The Topics to Focus On for the Midterm: IMPORTANT WARNING !! If You Don't Attend Each Class, You Wouldn't Know Which Lectures and Which Lecture Slides Were Covered in Class for The Midterm and the Final !! Characteristics and Architecture of Big Data Processing Systems/Applications, Cloud Computing Systems Universal Data Exchange Formats in Object Exchange Model(OEM): Three Common Big Data formats in Semi-Structured Model: Data Model, Syntax of Each Encoding Format, the Related Data Processing Techniques: DOM, XPath Characteristics of Semi-Structured Data Model, Main Differences between Semi-structed Data Model and Relational Data Model for Object Relation Mapping (ORM), Conversion among Relation(CSV), XML, and JSON Semi-Structured Database System -- MongoDB: CRUD Basic Operations, MongoDB Queries for Embedded Objects and Array, Aggregation Pipelining Comparison Between Semi-structured Database Server and Relational Database Server for Database Management Strategies. NOTE that ONE Page Note is NOT allowed to the Midterm !! Exam Formats: 6-7 Main Questions with 2-3 sub questions for a Small/Short Answer Types for Problem Solving Final Exam Info for 2026: Monday May 4 4:00p-6:00p In Each Lecture Note for the Final, Only the Slides (Subjects) Covered in Class Will Be in the Final See the Lecture Notes Links Only for the Final here (All the rest of links are disabled) //a href="CIS612_FinalLinksOnly.html"> the Lecture Notes Links Only for the Final One Page Note (Each Side) Is Allowed to The Final. It Shoule Be Hand-Written. Printed or Copy and Paste of Entire Letcure Notes Are NOT Allowed. Topics to Focus on for the Final (Tentative) Unstructured Data Processing, Inverted Index, How to Build Inverted Index, TF-IDF Based Ranking Algorithm with Cosine Similarity Text Cleaning Preprocessing, Basic Natural Language Processing (NLP) Methods in Data Pipeline, POS Tagging Problems of TF-IDF Based Ranking for Text Analysis Context Aware NLP Methods - POS Tagging, NER Tagging Design of Intelligent System with Data Pipelining of Big data processing Semi-Structured Query Model - XQuery Information Extraction Methods, How to Build a Knowledge Base from a Large Collection of Unstructured Texts for Question Answering System Semi-Supervised Learning - word2vec and Question-Answering System XQuery Using XPath Intro to RDF Concepts Graph Database Neo4j Cypher Query Basics Parallel Map Reduce Programming, Parallel Distributed Data Processing Execution Algorithm in MapReduce (the Execution Phases (Steps) of Map Reduce on HDFS), Data Execution Algorithm Steps in Map phase and Reduce phase, Architecture of HDFS -- Architecture of Parallel Distributed File System in General Architecture of Parallel Distributed Platform (Amazon EC2 Cloud Platform in General) MongoDB Queries: Aggregate Pipelining, Look Up Join Operator, Hive DDL to Build DW like Data Cubes, Mapping from HiveQL to Map Reduce Jobs, Pig Latin Data Pipelining Execution Steps for General Analytic Queries Join Algorithms on a Parallel Distributed System in Map Reduce Data Models of NoSQL Systems: Hive, PigLatin, MongoDB, Google Big Table/HBase (Not for This Semester) Transaction (Random Reads/Writes) for Concurrency Control of NoSQL Systems: Hive, MongoDB, Google Big Table/HBase Spark Data Processing Architecture Kafka Data Processing Models/Architecture (Not For This Semester) One Page Hand Written Note (Each Side) is Allowed to the Final ! Printed Copy and Pasted Lecture Notes Are NOT Allowed as One Page Note |
|
Projects on Big Data Processing, Training Large Language Models with Big data, Building an AI Application with Big Data Analytics Project Guideline and Important dates ! //a href="TaskListTimelineBigDataProject.pdf"> Project Task Timeline Important Dates for Project: (Tentative) Proposal Submission by April 12 ! Status Report Submission by April 21 ! Presentation Starts on April 27 (Mon) and 29 (Wed) ! Final Project Report by Saturday May 1 ! Important Notes for Final Project: 1. You Can Do One Person Group Project. 2. 3-4 Person Group is Allowed with a Bigger Scale Project 3. If You Need to Find Group Members, You Can Use the Blackboard Class Email to Send to the Class, Some will respond to you if they are looking for a group member. 4. Implementing Your Final Project with Clientside and Serverside of a Web Application with a Web Based User Interface Is NOT Required. (It will Be Counted More If AI/Big Data Project is impleneted as a Web Application (Created as a Website with a Webserver). 5. You can Change Your Project Plan/Proposal or the Details of Proposed Project even after Your Proposal has been submitted until the Deadline of the Status Report. 6. For Those Who Have Aleady Taken CIS660, it is Required to Complete a Full Project with Bid Data Processing and Training ML to Have Predictive Modeling 1 - 2 Person Group Project Are Allowed for the Small Class Size (~25 students); 1 - 3 Person Group Project Are Allowed for the Big Class Size (> 35 students); 1 - 4 Person Group Project Are Allowed for the Very Big Class Size (> 40 students) If 3-4 Person Group is Allowed, Make Sure to Make the Project in a Bigger Scale If 3-4 Person Project Group Will Be Allowed As Long As the Project Scale is Big Enough for 3-4 Persons PhD Student Should Do One-Person Project (Not in a Group) Final Group Project Specification and Instructions: Project Submission Instructions: One Submission Per Group Required. Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week ! Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx ! If your data file is too big to upload, Submit your zip file with your Data file either on the google drive or One Drive and send the link. (Send me email for access permission for this link !) Submit a Zip file that includes: - All of your Group Presentation Slides (Must be .ppt) - Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files on Blackboard by the end of Friday of your presentation week. - Include the Problems/Error Encountered and Your Resolutions in Your Report - Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files. - The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. - If Your Group Report/Presentation Don't Show/Include Any of the Required Contents in Your Report and Presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes that were downloaded from the Internet. Grading Criteria of Project Complexity Based on: - Big Data Size, - Complexity of Big Data Processing/Transformation Methods - Complex Big Data Processing Methods or Analytic Scoring/Ranking Methods - Superior Knowledgebase Design Such as Grapgh or Property Based Complex Knowdlege Representation Design (Naive Structured CSV/TSV Data or Already Processed File Downloaded from Kaggle sites Will Not Be Considered as Big data) Preferred Big Data Project Platforms: They are not required but Will Be considered more ! 1. The Platform/Systems with Any Parallel Database Server with MapReduce and Hadoop Distributed File Sysrem Or 2. Real-time Big Data Analytic Processing with Spark Some Suggested Real life Big Data Sets for Group Projects //a href="https://www.yelp.com/dataset"> JSON Data sets at Yelp Site Social Media Twitter Server Generated Log Stream Data in JSON: //a href="CIS612.zip">Social Media Twitter Data in JSON: 2020 Presidential Election Candidate Trump and Biden Data Sets //a href="RawJsonTwitterData.zip">Social Media Twitter Data in JSON: farmers protest twitter data set //a href="https://medlineplus.gov/encyclopedia.html"> Medline Plus for Medical Encyclopedia //a href="https://dumps.wikimedia.org/"> 7.2 Million Wiki Page Dump either in HTML or XML //a href="https://dumps.wikimedia.org/mirrors.html"> 7.2 million wiki webpage dump in XML //a href="https://ftp.ncbi.nlm.nih.gov/pubmed/updatefiles/"> 32 million PubMed Collection of Biomedical Research Paper Abstracts in XML Any document collection of clinical notes or doctors’ notes //a href="https://mimic.mit.edu/"> MIMIC: Medical Patient Data to get Doctors’ Text Notes The IMDB Data links in the Project List are gone. Check here for Movie Review Data Set. //a href="https://www.kaggle.com/rounakbanik/the-movies-dataset#movies_metadata.csv"> Movie Review Data //a href="https://developers.google.com/youtube/reporting/"> Youtube Analytic Site Examples of Good Final Big Data Projects For Those Who Have Already Taken CIS660, The Final Group Project Should Include Fully Analytic Processing Suggested Projects: For Text Analytics like Sentiment Analysis or Opinion Analysis: NLP Techniques - POS, NER Tagging, Bi-Gram Handling Are Required for Preprocessing. For Document Categorization: by Constructing TF-IDF Vectorization. Inverted Index Building Will Be Plus but Optional. Building Word2Vec Embeddings for a Collection of Documents/Webpages with Training Set Generation in Skip Gram Model For Other Types of Projects, the Proposal is Required to Be Approved to Meet the Complexity of Final Project See the Project Page for Common Project Ideas or Examples of Big Data Projects //a href="https://eecs.csuohio.edu/~sschung/cis612/CIS612Project.html"> Group Project *********************** Project List is posted. //a href="BigDataProjectList_CIS612.pdf"> Project List IMPORTANT Submissions for Group Project Task 1: Group Project Proposal (Plan) Submit minimum 3 Page Group Proposal on Blackboard with List of Group Memebers Group Project Proposal (Plan) Should Include Brief Descriptions on: 1. Description of Big Data in Size and Format, Data Collection Plan, 2. Goal of Your Big Data Analytic Project with What Kind of Intelligent Analytic Funtionality, Features of Your AI/Big Data Analytic Application 3. Big Data Processing Plan, Methods 4. Investigate on Platform/Systems/Tools/APIs to Use Task 2: Group Project Status Report Your Project Status Report Should Show the Following Tasks Done: 1. Your Big Data Collection Done 2. Platform Setting/System Configuration Procedure (if it is new) 3. Design of Big Data Processing Pipeline, Data Transformation Methods 4.Your KnowledgeBase Structure/ Database Design 5. Your Applucation/Analytic Server Codes, Processed Data/Database Contents in Progress, Any Intermediate Outputs in Progress Task 3: Project Presentation Project Presentation Starts from the Last Week of the Class of the Semester Project Presentation Schedule Will Be Sent To Your CSU Email for Sign Up One week Before the Presentation Read the Instructions of Project Presentation Here ! Group Project Presentation Should Include: 1. Data Description, Data Size, Data Collection Method 2. Goal of Your Intelligent Big Data Analytic Application (AI) 3. Platform Setting/System Configuration Procedures 4. System Design (Architecture) of Your AI Application in Detail 5. Raw Big Data Preprocessing Methods and Intermediate Results 6. Design of Big Data Processing Pipeline, Data Transformation Methods 7. Description of Your KnowledgeBase Structure/Database Design and Real-time Question to Query Processing for Your AI Application. Show the Contents 8. Your ML Algorithms if any, Ranking Algorithm if any for Labelling if any 9. Your ML Models if Any and Evaluation (Accuracy) Results and Visualization of the Result, Data Marix (Structure) for Analysis/Evaluation and Visulation of the Analysis Results (for example, correlation matrix or similarity matrix) if Any 10. The Problems/Errors Encountered and Your Resolutions 11. System Demo Submit Final Project Report and Presentation By the End of Friday of The Last Class Week After Your Presentation! One Submission Per Group Required. Submit a Zip file that includes: 1. All of your presentation slides (both in .pptx) and 2. Your Group Final Project Report (in doc) The Final Project Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. The Report Should Include Platform/System Set up the Set Up Procedure /Configuration Detail of Your Platform/System/Packages, Executions Steps, all the Source Codes, Scripts, all the intermediate outputs, and final output files. Include the Problems/Error Encountered and Your Resolutions in Your Report If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web. Group Project Presentation Should Include: 1. Data Description, Data Size, Data Collection Method 2. Platform Setting/System Configuration Procedures 3. System Design (Architecture) of Your AI Application or Data Analytic Goal in Detail 4. Raw Big Data Preprocessing Methods and Intermediate Results 5. Design of Big Data Processing Pipeline, Data Transformation Methods 6. Description of Your KnowledgeBase Structure/Database Design. And Show the Contents 7. Ranking Algorithm, Data Matrix (Structures) if any for Evaluation 8. The Problems/Errors Encountered and Your Resolutions 9. System Demo or Evaluation Results and Visualization of the Result Project Examples of Vectorization of each document in a Training set for a Machine Learning Classifier: //a href="https://eecs.csuohio.edu/~sschung/CIS660/ProjectExample_SentimentAnalysis_YelpReview.pdf"> Project Example I: Sentiment Analysis of Yelp Business Review with Machine Learning //a href="https://eecs.csuohio.edu/~sschung/CIS660/ProjectExample_SentimentAnalysis_YelpReview_Feng.pdf"> Project Example II: Sentiment Analysis of Yelp Business Review with Machine Learning Best Group Projects: Selected Best Projects Will be Posted Here ! Big Data and Data Science Projects: More To Come Here ! Group Project Data Sources: You can choose to work on these data sets for your group project Final Project Submission Instructions: Submit Group Project Presentation and Final Report in a Zip File By the End of Friday of Your Presentation Week ! Remember you have to include the source file of your Project Report in doc and Presentation slides in pptx ! If your data file is too big to upload, Submit your zip file with your Data file on your google drive or One Drive and Send email to me and TA to share ! Submit a Zip file on Blackboard by the end of Friday of your presentation week. One Submission Per Group Required. Your Project Zip File Should includes: 1) All of your presentation slides (both in .ppt and .pdf) and 2) Your Group Final Project Report (in doc) with Platform/System Set up Procedures/Instructions, Executions Steps, all the source codes, scripts, all the intermediate outputs, and final output files 3)Include the Problems/Error Encountered and Your Resolutions in Your Report Your Final Project Report Should Include the Set Up Procedure /Configuration Detail of Your Platform/System/Packages as well as Source Codes and Intermediate Results in files. The Report Should Explain Each Step of Your Project Tasks with the Screen Captures and Results. IMPORTANT NOTE !!! If you don't show/include any of the required contents in your report and presentation, I will ASSUME that your group submitted a Copy of Somebody's Github Codes your group downloaded from the Web. |
|
Basic Python Tutorials : Guides for Important Platform Set up with Python, XPath, Beautiful Soup, and MySQL: //a href="CIS492_593_DSA469_Lab1_Environment_Setup.pdf"> Platforms Set up Guides for Big Data Labs (by TA Dennis Risch) ******************************* See More Guides in the Lab1 Section (Far Below) Python Data Science Platforms: //a href="https://www.scipy.org/install.html"> Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system) • Anaconda //a href="https://www.anaconda.com/open-source"> Anaconda Open Source Site See Fundamental Section for List of Data Science Platforms //a href="https://docs.anaconda.com/anaconda/navigator/tutorials/"> Anaconda Tutorials • Python Anaconda Tutorial Sites //a href="https://data-flair.training/blogs/python-anaconda-tutorial/"> Anaconda Tutorial Site //a href="https://www.edureka.co/blog/python-anaconda-tutorial/"> Anaconda Tutorial Site • PyTorch //a href="https://pytorch.org/"> PyTorch Site (It can be integrated from Anaconda as well) Google Colab: //a href="https://www.dataquest.io/blog/jupyter-notebook-tutorial/"> Basic Guide for Python tools for Data Science: jupyter-notebook //a href="https://www.dataquest.io/blog/advanced-jupyter-notebooks-tutorial/"> More Basic Guide for Python tools for Data Science: jupyter-notebook • Python Scikit Learn for Common Data Science Tasks Text Preprocessing (Natural Language Processing) Library in Python SpaCy: //a href="https://spacy.io/api"> Liquistic Modules in Python SpaCy //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing • Python sklearn.cluster //a href="https://scikit-learn.org/stable/modules/clustering.html"> Python Sklearn Clustering Python IDE Deduggers: //a href="Pycharm_Debugger"> Basic Guide for Python Debugger Pycharm //a href="https://www.spyder-ide.org/"> Python IDE Spyder //a href="http://docs.spyder-ide.org/current/panes/debugging.html"> Python Debugger Spyder Basic R Tutorials : //a href="IndependentStudyCIS611Final Report.pdf"> Special Online Study Guides on Basics on Data Warehouse/OLAP, Data Analytics, Big Data in Independent Study Independent Study with Nick White (Now in FaceBook and The First Prize Winner of 2016 Senior Project) Useful Machine Learning Tutorial Sites: Keras for Image Processing/Text Processing with Deep Learning: //a href="https://keras.io/guides/"> Keras Machine Learning For Your Own Advanced Study Lab Submission Instructions: 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof. 2. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System. 3. If You did Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front Page of Your Report in Bigger and Bold Font ! Useful Lab Helpers: Useful Tools: //a href="https://swagger.io/"> Swagger API Tool for REST API Developments Useful Big Data Analytic Tools Choose your System/Tool/Platform to Set Up and Get Used to: Machine Learning in Python: Natural Language Processing (NLP) for Text Preprocessing/Big Data Analytics Library in Python SpaCy: //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing Lab Assignments: Instructions for Lab Submission: The Lab Submission Link and the Deadline of Each Lab Will Be Posted on the Class BlackBoard ! Wait For the Lab Submission Links Are Created on Blackboard for Each Lab You Have to Start Working on Labs Before the Submission Link Are Created on Blackboard for Each Lab Submission. Always Follow the Deadline of Each Lab Assigned on the Class Blackboard. The Deadlines mentioned on the Class Webpage Are Tentatively Scheduled at the Beginning of Each Semester. Please Identify Your Course When You Ask Me in Email ! Each of your screen capture must show your results returned by your systems and your database servers to prove that you have done the lab correctly !! 1. Submit your Zip file that includes your report in .doc file (that expains all the platform set up, the execution steps, and copy of each source code files ) and all the Source files, and output files on Blackboard for a timestamp and as a proof. 2. IMPORTANT NOTE: Your Screen Captures in Your Lab Report Should Show Your Own System and Your Database Server Name to Prove That Your Lab Was Done In Your System. 3. If You did Extra Credit Part, Mention about What Part is Done for Extra Credit at the Front Page of Your Report in Bigger and Bold Font ! Lab0: Learning Python -- Due by the End of the Second Friday of the Semester Python Data Science Platforms: See More Python Platforms Above or Lab1 Set Up Guides Below to Choose for Data Science //a href="https://www.scipy.org/install.html"> Installation Guide for Scientific Python tools for Data Science with pip (inbuilt package management system) Installation Guides for MySql Server: //a href="https://www.mysql.com/downloads/"> MySql Download //a href="https://dev.mysql.com/doc/refman/5.7/en/installing.html"> MySql Download and Installation //a href="https://dev.mysql.com/doc/refman/8.2/en/tutorial.html"> MySql Tutorial //a href="https://dev.mysql.com/doc/refman/5.7/en/creating-database.html"> How to Create MySQL Database If You Want to Use MS SQL Server, Installation Guides for MS SQL Server: See the Announcement Section of CIS430/530 Database Systems and Processing below for Account Creation for Microsoft Azure site for Free Download of MS Visual Studio and MS SQL Sever. //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Announcement"> How to Create MS Azure Potal Site Account See the Lab Section of CIS430/530 for Installation Instruction of MS Visual Studio and MS SQL Sever. //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> How to Download and Install MS SQL Server XML/XHTML/JSON Syntax Validators: //a href="https://www.infoplease.com/homework-help/history/collected-state-union-addresses-us-presidents"> Infoplease site of State Union Addresses of US Presidents //a href="https://www.infoplease.com/homework-help/us-documents/state-union-address-john-adams-december-3-1799"> Correct page of Address of John Adams December 3 1799 //a href="SimplifiedInfoUnionAddress.htm"> The Simplied html file of the Info site (You may use this simplified html file to inspect the element structure to extract the required info for Lab1) Important Notes: Infoplease site of State Union Addresses of US Presidents (As of 2025, This site blocks any http request from unknown clients (you)) You can either using the workaround below to avoid blocking or use another site for State Union Addresses of US Presidents Another site for State Union Addresses of US Presidents that Does NOT block http request Implementation Related Important Notes: If there is a Link that Does NOT Have Any Web Page Contents, Add NULL Values for the corresponding Columns for the link Do not Assume that every sites has an identical URL format. This is semi-structured data. Nothing is regular in Big data. For the irregular parts, use regular expressions or xpath as neccessary. For Those Who Took CIS593 Big Data: Lab3 on Information Extraction from Customer Reviews on the Social Media Websites Using DOM and XPath Visit Yelp Cleveland site below to Extract each customer review in each restaurant page. //a href="https://www.yelp.com/search?find_desc=Restaurants&find_loc=Cleveland%2C+OH">Yelp Site for Best Restaurants in Cleveland, OH For example, Extract each customer review in the restaurant page to store them in a SQL database table //a href="https://www.yelp.com/biz/lj-shanghai-cleveland?osq=Restaurants">Website for a Restaurant with Customer Reviews Create a table with Restaurant name, Location(Address), Reviewer Name, The number of stars, Review Text Make it Automated for the Table Creation from your output file (CSV file) in your Database Server ! General Approaches for Webpage Processing There are two ways to do Information Extraction from Webpages. Method 1 with HTML DOM and XPATH is REQUIRED for Lab1 !!! Method 1: Webpage as Semi-Structured HTML DOM Tree Using XPATH (More Scores Given!) Method 2: Webpage as Unstructured Text Using Parsing API like Beautiful Soup Set Up Guides for Important Platform with Python, XPath, Beautiful Soup, and MySQL:****************************** You Need To Set UP DOM and XPATH for Any Server-side Scripts/Languages for Applications such as Lab1. //a href="CIS492_593_DSA469_Lab1_Environment_Setup.pdf"> Platforms Set up Guides with Python, MySql, XPath, PyCharm Debugging IDE for Big Data Labs (by TA Dennis Risch) ******************************* More on Setup Guides: //a href="https://pip.pypa.io/en/stable/quickstart/"> pip Python package Installation Guide //a href="Setup Guide for Anaconda Python and Jupyter Notebook.pdf"> Setup Guide for Anaconda Python Framework and Jupyter Notebook IDE (by TA Hemal Paneliya) //a href="Setup Guide for Anaconda Python and SpiderIDE.pdf"> Setup Guide for Anaconda Python Framework and Spider IDE and Debugger (by TA Durga Dasepalli) pyodbc Set Up Guide With MS SQL Server: //a href="HowtoConnectPythonJupyterToSQLServer_pyodbc.pdf"> How to Set Up Python Jupyter Notebook to Connect to a Microsoft SQL Server using pyodbc**************** Platform Set Up Guide on Mac: //a href="Installation guide for Python lxml in Pycharm IDE.pdf"> Installation Guide on Mac for lxml with pip Python package for DOM with XPath and pycharm IDE for debugging on Mac / More on Python IDE Deduggers: //a href="Pycharm_Debugger"> Basic Guide for Python Debugger Pycharm //a href="https://www.spyder-ide.org/"> Python IDE Spyder //a href="http://docs.spyder-ide.org/current/panes/debugging.html"> Python Debugger Spyder Database Server Set up is needed for the Labs. You can use any SQL Server -- MySQL, MS SQL Server, or any Database Server See Lab0 Section Above for more instructions or See the step-by-step installation guides in CIS430/530 Lab Section Below //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> CIS430/530 Lab Section //a href="Setup_MSSQLSever.pdf"> How to Download and Install MS SQL Server Note that the Set Up Guides Above are for Server Side Applications such as Lab1 in Python or Java. You don't Need to Do Any Extra Set Up to Use XPATH in a JavaScript in a Client side Codes for HTML DOM Processing which will be executed by your Web browser. All the Recent Web Browsers Have XPATH Features in their Debugger by Default. IMPORATNT NOTE !!!! The tutorials and examples in the class Lab Section are to guide and teach the methods and techniques. Please note that they are NOT to provide you with precise coding solutions that would run without checking/debugging the changes in codes; Especially since these areas are fast changing and there are so many variations in the platforms/languages. YOU ARE RESPONSIBLE for DEBUGGING YOUR CODES !! Lab1 Implementation Guides:******************************* //a href="Lab1Guides_Links for XPath_Pyodbc_BLOBDataType.pdf"> Lab1 Guides: What to Need to Know in the Lab1 Section to Do Lab1 ************************* Note that the Code Examples of the Lab1 Guides below Do NOT Contain the Complete Codes to Be Executed. These are ONLY for the guides for Lab1. The Environment Configuration and the Versions of APIs Varies. Do Not Copy the Entire Codes Blindly to Do Your Lab1 since it will not work depending on the version of your python and setting up for DOM and XPath. How to Write Xpath in Demo in Webbrowser: //a href="Lab1_Xpath_Syntax_WebpageDemo.pdf"> Example of Xpath Execution in Webbrowser ************************* //a href="Lab1_Xpath_WebpageDemo2.pdf"> Code Examples with DOM and XPATH: //a href="CIS593_Lab1_Guide_InfoExtraction_WebPageEmery_Partial.pdf"> Example of Lab1: General Guide in Python with DOM and XPath, and pyodbc for database operations for Information Extraction ************************* //a href="CIS593_Lab1_Guide_Web_InformationExtractionDOMXpath_A.pdf"> Example of Lab1: Guide in Python with DOM and XPath in lxml for Information Extraction ************************* //a href="CIS469_569_Lab1_GuideWebScrappingDOMXpathwithPHP.pdf"> Example of Lab1: Guide in PHP with DOM and XPath for Webscapping //a href="CIS593_Lab1_WebscrapDOMXpathGuideLinux_CFord.pdf"> Example of Lab1: Guide in Python on Linux for Webscapping Code Examples with Beautiful Soup Text Processing API: //a href="CIS469_569_Lab1_GuideWebScrapping.pdf"> Example of Lab1: Guide in Python with Beautiful Soup for Information Extraction DOM with XPATH Set Up in Any Script/Programming Languages for Server Side Applications Note that You don't Need to Do Any Extra Set Up to Use XPATH in your Serverside JavaScript in NodeJS. All the Labs of CIS612 Are Server-side data processing in a server-side script/programming language with file I/O and DOM parser with Xpath. They are NOT a client side JavaScript executed by your Web browser. For the Labs to Write an Application Server in this course, You Need To Add DOM Parser related APIs for Your Scripts/Programming Languages To Do the following steps: to Call a DOM Parser API to Build a DOM Tree and to Use XPATH Methods for Retrieval to Extract information You need as in the Examples below. Any Modern Script/Programming Languages Have the DOM Parser and XPath Features. You Need to Set up to Use. //a href="http://docs.python-guide.org/en/latest/scenarios/scrape/">Setting Up lxml parser for DOM with XPath for Python ***************************** //a href="NodeJSSetupXPathXMLInJS_2.pdf"> Node.JS XPath Setup Guide and XML Parser in JavaScript based Node JS //a href="https://github.com/goto100/xpath"> XPath Setup Guide for JavaScript //a href="https://www.npmjs.com/package/xpath"> XPath Setup with npm in JavaScript for Node JS Set Up DOM with XPath for HTML and XML: //a href="HowtoDebugHTMLJAVAScriptXPath.pdf"> How to Debug HTML with JavaScript, DOM, XPath //a href="https://search.yahoo.com/search?fr=mcafee&type=E211US1289G0&p=how+to+inspect+element+in+chrome"> Learn How to Inspect Element in Chrome to Debug HTML with for DOM and XPath JavaScript: //a href="https://www.w3schools.com/js/js_debugging.asp"> JavaScript Debugger //a href="https://www.w3schools.com/xml/xpath_intro.asp"> XPath Examples in Javascript //a href="https://developer.mozilla.org/en-US/docs/Introduction_to_using_XPath_in_JavaScript">MDN Site for DOM with XPath in JavaScript //a href="https://github.com/goto100/xpath"> DOM with XPath Setup Guide for JavaScript //a href="https://github.com/goto100/xpath/blob/master/docs/xpath%20methods.md"> XPath Methods //a href="NodeJSSetupXPathXMLInJS.pdf"> Node.JS XPath Setup Guide //a href="https://www.npmjs.com/package/xpath"> npm Node.JS XPath Setup Guide //a href="https://chrome.google.com/webstore/detail/xpath-helper/hgimnogjllphhhkhlmebbmlgjoejdpjl"> Google Crome XPath Helper Setup Python: //a href="http://docs.python-guide.org/en/latest/scenarios/scrape/">Setting Up DOM with XPath for Python //a href="https://www.crummy.com/software/BeautifulSoup/bs4/doc/"> BeautifulSoup Documentation for HTML, XML for Python //a href="http://web.stanford.edu/~zlotnick/TextAsData/Web_Scraping_with_Beautiful_Soup.html"> Webscapping with DOM, XPath for Python with BeautifulSoup (From the Stanford Class) //a href="https://github.com/helloitsim/InstAnalytics/blob/fdec9d27896f4b6e61acd9dc56134ad556144057/InstAnalytics.py"> Python XPath Guide Automatic Table Creation in a SQL Server After installing a SQL server (See the step-by-step installation guides in //a href="https://eecs.csuohio.edu/~sschung/cis430/CIS430IDS.html#Lab"> CIS430/530 Lab Section Add the codes in pyODBC as below for a Connection for the Anaconda Python Framework to Connect to Your SQL Database Server For a connection with SQL Server with servername and database name using pyodbc in Python: conn = pyodbc.connect('Driver={SQL Server};Server=YOUR_SQL_SERVERNAME\SQLEXPRESS;Database=YOUR_DATABASE_NAME;Trusted_Connection=yes;') pyodbc(Open Database Connectivity) to Connect a MySql Server in Python: //a href="https://docs.devart.com/odbc/mysql/python.htm"> How to Connect to MySql Server with pyodbc in Python //a href="https://stackoverflow.com/questions/3982174/pyodbc-and-mysql"> Stackoverflow site for pyodbc: Examples to Connect to MySQL pyodbc(Open Database Connectivity) to Connect MS SQL Server in Python: //a href="https://docs.devart.com/odbc/sqlserver/python.htm"> How to Connect to MS SQL Server in Python //a href="https://datatofish.com/how-to-connect-python-to-sql-server-using-pyodbc/"> How to Connect to MS SQL Server in Python //a href="https://stackoverflow.com/questions/33725862/connecting-to-microsoft-sql-server-using-python"> pyodbc: Examples of ODBC in Python //a href="https://docs.microsoft.com/en-us/sql/connect/python/pyodbc/step-1-configure-development-environment-for-pyodbc-python-development?view=sql-server-ver15"> Installation pyodbc to Connect to MS SQL Server in Python //a href="https://docs.microsoft.com/en-us/sql/connect/python/pyodbc/step-3-proof-of-concept-connecting-to-sql-using-pyodbc?view=sql-server-ver15"> Stackoverflow site for pyodbc to Connect to MS SQL Server in Python //a href="http://www.sqlines.com/oracle/datatypes/clob"> Column Data Types to store a Large Text data to Create a Table with in MS SQL Server or other database server How to Create a Table in a SQL Server from CSV/TSV Text files //a href="https://docs.microsoft.com/en-us/sql/t-sql/statements/bulk-insert-transact-sql?view=sql-server-2017"> How to Create a Table from a File with Bulk Insert with MS SQL Server //a href="https://stackoverflow.com/questions/14330314/bulk-insert-in-mysql"> How to Create a Table from a file with Bulk Insert in MySQL Server //a href="FullBulkImportAWIDSQLServer.sql"> Example of a Script to Create Multiple Tables from different files with Bulk Insert in MS SQL Server Earlier Features to Handle Big Data in Relational Database Server Advanced Data Types: BLOB (Binary Large Object) or Text/CLOB (Character Large Object) in MySQL or MS SQL Server How to Create and Insert to a Table with a Column of Large Text Data or Image Data in a Relational Database Server: //a href="https://www.mysqltutorial.org/mysql-text/"> Text Data Type in MySQL //a href="https://www.tutorialspoint.com/What-is-TEXT-data-type-in-MySQL"> What is TEXT data type in MySQL //a href="https://www.tutorialspoint.com/what-is-the-difference-between-blob-and-clob-datatypes "> Difference between blob and clob datatypes //a href="https://dev.mysql.com/doc/refman/8.0/en/blob.html "> Blob Data type in MySQL //a href="https://stackoverflow.com/questions/10729824/how-to-insert-blob-and-clob-files-in-mysql"> how to insert blob and clob from files in mysql //a href="https://stackoverflow.com/questions/752277/cannot-insert-string-into-mysql-text-column "> how to insert text column in mysql For Those Who Want to Use CLR Table Function to Create a Table in SQL Server -- This is NOT For CIS492/593 //a href="HowtoSETUPUDTASPNET.pdf"> How to Set Up ASP.NET with SQL Server //a href="HowToDebugSQCLRNOExamples.pdf"> How to Debug CLR UDF, CLR UDT, CLR TVF //a href="https://datatofish.com/how-to-connect-python-to-sql-server-using-pyodbc/"> How to Connect to MS SQL Server in Python //a href="https://stackoverflow.com/questions/33725862/connecting-to-microsoft-sql-server-using-python"> pyodbc: Examples of ODBC in Python Automatic Table Creation in a SQL Server pyodbc(Open Database Connectivity) to Connect MS SQL Server in Python: //a href="https://docs.microsoft.com/en-us/sql/connect/python/pyodbc/step-1-configure-development-environment-for-pyodbc-python-development?view=sql-server-ver15"> Installation pyodbc to Connect to MS SQL Server in Python //a href="https://docs.microsoft.com/en-us/sql/connect/python/pyodbc/step-3-proof-of-concept-connecting-to-sql-using-pyodbc?view=sql-server-ver15"> pyodbc to Connect to MS SQL Server in Python //a href="https://datatofish.com/how-to-connect-python-to-sql-server-using-pyodbc/"> How to Connect to MS SQL Server in Python //a href="https://stackoverflow.com/questions/33725862/connecting-to-microsoft-sql-server-using-python"> pyodbc: Examples of ODBC in Python How to Create a Table in a SQL Server from CSV/TSV file //a href="https://docs.microsoft.com/en-us/sql/t-sql/statements/bulk-insert-transact-sql?view=sql-server-2017"> How to Create a Table with Bulk Insert with MS SQL Server //a href="FullBulkImportAWIDSQLServer.sql"> Example of a Script to Create Multiple Tables with Bulk Insert in MS SQL Server Lab3_2: //a href="CIS612_Lab3_XML_2024.pdf"> Lab3_2 on XML Processing with DOM and XPath for ORM //a href="bibXMLInputNoDup.pdf"> Lab Assignment 3_2 - XML Input File //a href="CIS593_Lab2_JSON_FAQs.pdf"> General FAQs on Semi-Structured Data Transformation to a Relational Scheme Check the validity of XML Input. Add a proper XML doctype heading like in your XML input document and Validate your XML input and correct it if needed. //a href="XMLDomParserExample.pdf"> LabAssignment 3 - DOM XML Parser Example Codes //a href="Lab3OutputExampleLawrence.pdf"> Lab3 Output Examples of Conversion Between XML and SQL Table //a href="Stored ProcedureCreateTablesfromCSV.pdf"> Example of Stored procedure to Create SQL Tables from a Text file JAVA based XML DOM Parser to download to set up if you are using JAVA for this lab. //a href="http://www.w3.org/DOM/"> DOM XML Parser //a href="http://www.saxproject.org/"> SAX XML Parser //a href="ExamplesXMLToTable.pdf"> Examples Conversion Between XML and Table //a href="FAQOnLab3_2019.pdf"> FAQ on Lab3_XML and Lab3_JSON Lab 3_3 on ORM Mapping with Semi-structured JSON Data Processing Notes and Corrections: Note that you need to transform business.json file only for Lab3_3, not all of 5 JSON files from the Yelp site. Yelp Business Data Set for Lab3_3: Data Sets for Lab3_3: The given OneBusiness.json file and business100.json file in JSONData.zip are the corrected files. Yelp Full Data Set Also Available in the Big data Lab below: //a href="YelpDataSets.zip"> Zip file for 5 JSON files from Yelp data Challenge 2017 or Download directly from the Yelp site below for 2020 data sets! Note: If You are having a corrupted zip file error for 2017 Yelp Data Set, use 7-Zip to unzip. You can directly download the most recent data sets from Yelp site below Yelp Data Set and Documentation for JSON File Structures Note: Those Full Data Sets May Not Be in a Valid JSON Format. (Always expect this in a real world project) You Have to Use 7-Zip to unzip the zip file. Some other compression software might not be able to unzip correctly. The full JSON data files from the Yelp site might have a few incorrect JSON syntax detected in the data file or invalid line feed, which is common. Correct them before processing if detected. See FAQs for how to correct invalid JSON data Invalid JSON format handling: If there is any invalid data format is found in the input file, you can change it to the correct JSON format. For example, $$ in the “Price Range” key value pair in your input json file, the value $$ is not in quotes and this will cause to fail. Suggested Solutions: Replace $$ with 2 (meaning the price level is 2 in scale 1 - 5) in the file and try to parse the corrected file in your program. If , is missing between objects, add it. If you have invalid JSON file problem, the JSON file might need to be corrected through the JSON validator to detect the errors first then correct the JSON syntax to be imported. The data files in the Zip file are all coming from the yelp site. They sometime have the incorrect JSON syntax, which need to be corrected to make it work. See the example of how to correct an invalid json file below. Scroll down for the part. //a href="https://eecs.csuohio.edu/~sschung/CIS593/CIS593_Lab2_JSON_FAQs.pdf"> FAQs for JSON Processing in General More FAQs on Lab3_3 Q: Can we use the 2019 dataset (JSON) from yelp website for our Lab 3_3? Or we should only use 2017 JSON data? A: Either data set is ok. Some people have a Zip error with the 2017 data set. If then use the 2019 data set from the Yelp site. Lab 3_4 on Semi-structured JSON Data Processing with MongoDB and ORM Creating Semi-Structured Database: MongoDB Collection in a MongoDB Server from business.json and review.json files from the Yelp site No //a href="CIS612_Lab3_4_Yelp_MongoDB_BusinessDataReviewData_Join_ORM.pdf"> Lab3_4 on Creating Collections from JSON data files of Yelp Business and Review Data (NOT For This Semester !) //a href="CIS612_Lab3_3_JSON_Collections_Yelp_MongDB_BusinessOnly.pdf"> Data Sets for Lab3_3 and Lab3_4: The given OneBusiness.json file and business100.json file in JSONData.zip are the corrected files. Yelp Full Data Set Also Available below: //a href="YelpDataSets.zip"> Zip file for 5 JSON files from Yelp data Challenge 2017 or Download directly from the Yelp site below for 2020 data sets! Notes: - There are differences between 2017 Business Data and 2020 Business Data. Either Data Set is OK for Lab3_3 and Lab3_4. - If You are having a corrupted zip file error for 2017 Yelp Data Set, use 7-Zip to unzip. You can directly download the most recent data sets from Yelp site below Yelp Data Set and Documentation for JSON File Structures Lab 3_4: NoSQL Database System MongoDB with Data Aggregate Pipelining for Information Extraction and Data Analysis //a href="CIS612_Lab3_4_JSON_Collections_Yelp_MongDB.pdf"> Lab3_4 on NoSQL Database System MongoDB with Data Pipelining for JSON data processing //a href="MongoDBQueryExamples.pdf"> Example Runs of Mongo DB Queries Create Semi-Structured Database in MongoDB for two JSON data files: Business.json and Review.json from the Yelp data site 1. Import Business.json and Review.json Data files from Yelp site into MongoDB 2. Retrieve Information Using MongoDB Aggregation Pipelining for Information Extraction Note that: You have to use a full business.json data and a full review.json data file. If you use 100 business data only, you will get an empty query result. If Q2_1 or Q2_2 returns an EMPTY result, Change the Filtering Condition review_count Until You Get the Reasonably Good Number of Results. For example, To avoid the empty query results, the filtering conditions have been changed to review_count > 5 and star <= 2. For Q2_1: the review_count > 5 For Q2_2: the review_count > 5 Note for MongoDB Join for Part2: If You are using 100 business data for the join opeartion, there are no matching business id in the business ids from the list of 100 businesses and the review collection, so join results were empty. The reasons for this could be either 100 business data is too small to get any join result with the review data or the review data set in the new data set from Yelp 2019 has been changed. They keep changing or updating the Review data sets, so there is no matching id in the 100 business data set in the class web page which was collected in 2017. Use the entire Business and Review json files in the new data sets from 2019 to have meaningful join results. Note that You have to process the entire data set in each json file, not just 100 Business data. Make Sure to Use All the Documents in business.json to create a Collection named business and another Collection named review for this Lab Yelp Full Data Set Available in the Big data Lab below: //a href="YelpDataSets.zip"> Zip file for 5 JSON files from Yelp data Challenge 2017 or Download directly from the Yelp site below for 2020 data sets! You can directly download the most recent data sets from Yelp site below Yelp Data Set and Documentation for JSON File Structures The given OneBusiness.json file and business100.json file in JSONData.zip given in Lab3_3 are the corrected files. Important Notes: If You are having a corrupted zip file error for 2017 Yelp Data Set, use 7-Zip to unzip. You Have to Use 7-Zip to unzip the zip file. Some other compression software might not be able to unzip correctly. Those Full Data Sets May Not Be in a Valid JSON Format. (Always expect this in a real world project) The full JSON data files from the Yelp site might have a few incorrect JSON syntax detected in the data file or invalid line feed, which is common. Correct them before processing if detected. See FAQs in LAB3_3 SECTION ABOVE for how to correct invalid JSON data If you have invalid JSON file problem, the JSON file might need to be corrected through the JSON validator to detect the errors first then correct the JSON syntax to be imported. The data files in the Zip file are all coming from the yelp site. They sometime have the incorrect JSON syntax, which need to be corrected to make it work. See the example of how to correct an invalid json file below. Scroll down for the part. //a href="https://eecs.csuohio.edu/~sschung/CIS593/CIS593_Lab2_JSON_FAQs.pdf"> FAQs for JSON Processing in General More FAQs on Lab3_4: Q: Can we use the 2019 dataset (JSON) from yelp website for our Lab 3_4? Or we should only use 2017 JSON data? And Do we need to create CSV for Q1 as well for the count result? A: Either data set is ok. Some people have a Zip error with the 2017 data set. If then use the 2019 data set from the Yelp site. From the results of Q2_1 and Q2_2 to extract the info (review_id, business_id, stars, review_text) to convert to CSV. Conversion is not required for the Q1 Result. Q: How many Extra Credit we will get if Lab3_4 is built as a full Rest API Application? A: It Will Be Counted 20% Extra Credit. //a href="https://docs.mongodb.com/drivers/pymongo/"> How to Install pymongo driver and Connect to MongoDB Server in Python application as a client See MongoDB Lecture Notes Section for MongoDB CRUD Queries and more details //a href="https://docs.mongodb.com/guides/server/import/"> How to Import Data file to MongoDB //a href="https://docs.mongodb.com/manual/reference/program/mongoimport/"> Mongo import //a href="MongoDBQueryExamples.pdf"> Sample Runs of Mongo DB Queries //a href="https://www.mongodb.com/compatibility/json-to-mongodb"> How to Import a json file to MongoDB //a href="https://www.mongodb.com/languages/python"> How to Use MongoDB in Python for CRUDE: DB/Collection Creation, Insert, and Querying //a href="https://www.geeksforgeeks.org/how-to-import-json-file-in-mongodb-using-python/"> How to Import a JSON file to MongoDB in Python //a href="https://stackoverflow.com/questions/70089317/how-to-do-a-word-count-in-mongodb"> Example of MongoDB Aggregation Pipeling for Word Count //a href="https://stackoverflow.com/questions/42941682/storing-json-data-into-a-variable-using-python-when-inserting-into-mongodb"> How to Save MongoDB Query Results into a variable //a href="https://stackoverflow.com/questions/20769621/saving-the-result-of-a-mongodb-query"> How to Save MongoDB Query Results Lab3_4: EXTRA CREDIT (30%) !!! Collecting Real Time Twitter Message Streams by Topics and Creating Semi-Structured Database in MongoDB for Retrieval If you choose Twitter Stream Data Collection to Do Lab3_4 instead of Yelp Data, it will be 30% Extra Credit ! Lab Specification: //a href="CIS612_Lab3_3_TwitterLogging_MongoDB_JSON_Transformation.pdf"> Lab3_4: Twitter Data Stream Collection, Collection in MongoDB, and JSON Transformation //a href="TwitterLoggingStructureJSON.pdf">Twitter Logging Structure in JSON Note some Tweets Don't have the Retweet part of info. Try to Collect Retweets Together. Note that: Lab3_4 Report to do Real Time Twitter Stream Data Collection and Use of MongoDB to Store and Maintain Them. The Collected Twitter Data Will Be Used for Sentiment Analysis as Your Final Project later //a href="Most Recent Update Twitter Free Account.pdf"> Most Recent Update on Twitter Free Account for Data Collection (By Collin Simpson) Twitter Data Sets: If You Can't Get Any Tweets Using Your Twitter Developer's Account, Use this Twitter Raw data set in this site: //a href="RawJsonTwitterData.zip"> Raw JSON Twitter Data (zip) -- farmers-protest-tweets-2021-2-4 //a href="coronatweets_11_53.zip"> Raw JSON Twitter Data (zip): coronatweets_11_53.zip //a href="2020ElectionTrumpBiden.zip"> Raw JSON Twitter Data (zip): Tweets on Trump and Biden (Collected One Month before the 2020 Presidential Election) //a href="https://www.kaggle.com/code/prathamsharma123/clean-raw-json-tweets-data"> Kaggle: Clean Raw JSON Tweets Data site Note That if TWITTER Real Stream Data Collection is not done, Extra Credit *will not be given ! For Twitter Stream Data Collection: NOTE that to Collect the Twitter Real Time Stream Data, you Need to Apply for Their Developer’s Account in the Twitter Developer's site You Need to Get a Permission to Get credentials for a token and keys. This process usually takes 3-4 days or one week for Twitter to respond. You HAVE TO Start Your Twitter Application ASAP. Do NOT Wait until the Last day. See the example project and sample codes for the step by step procedures for this. Read everything posted in the links below. For Your Twitter Developer Account Application: Apply a Twitter Developer's Account ASAP for Twitter Stream Data Collection ! Start ASAP Since it Will Take a Week to Obtain a Twitter Developer's Account !! DO NOT Choose Elevated, Advanced, or Academic Research Developer's Account to Apply. You Might NOT Get a Permission and It Will take More Time //a href="FAQs_Lab3_3_TwitterDataCollection.pdf"> FAQs to Get Twitter Developer's Account Twitter API 1.1 is depreciated and wont be available for new developers: //a href="https://twittercommunity.com/t/deprecation-announcement-removing-compliance-messages-from-statuses-filter-and-retiring-statuses-sample-from-the-twitter-api-v1-1/170500"> Twitter API 1.1 is depreciated The New Structure for the 2.0 API for tweets: //a href="https://developer.twitter.com/en/docs/twitter-api/data-dictionary/object-model/tweet"> the New Structure for the 2.0 API for tweets Note That the Recent Changes of the Twitter Site Seem to Affect Their Response Time to Process Applications for Developer's Account. Due to the Delay on the Twitter Site Response Time, You Need to Apply ASAP to Collect the Twitter Streaming Data in Time Twitter Stream Data Collection Collect at least 10,000 Tweets Talking about either One of the Following Topics of Your Choice. Add More Related Keywords As Needed to Your Chosen Topic to Collect As Many Related Tweets Possible: Suggested Topics: 1. Any New Major Movie or Product That Was Released Recently if Any (For example, IPhone - IPhone 12, IPhone Mini) Or 2. Any Major News (For example, 2024 US Presidential Election) Or 3. President or Any Person of Interest, or Any Two Candidates in an Election Or 4. Covid, Corona Virus, Covid-19 related topics Or 5. Any Topics of Your Interest as long as there are big enough to collect more than 10,000 Tweets Twitter Data Collection Setting Up: (The Twitter Data Collected will be used for Lab3 and Can Be Used For Your Final Project later) NOTE that to collect the Twitter stream data in real time, you need to apply for their developer’s account in the Twitter Deveoper's site and get a permission to get credentials for a token and keys. This process usually takes 3-4 days or one week for Twitter to respond. You HAVE TO start your application ASAP. Don’t wait until the last day. See the example project and sample codes for the step by step procedures for this. Read everything posted in the links below. Apply a Twitter Developer's Account ASAP for Twitter Stream Data Collection ! Start ASAP Since it Will Take a Week to Obtain a Twitter Developer's Account !! For Your Twitter Developer's Account Application, Choose the most Common Account type. Do not choose an Academic Reserach Account (You Will Be Asked More Questions). Do NOT Blindly Copy Those Sample Answers in the Class Webpage for Your Answers ! Rephrase/Modify Them in Your Words For Your Case. //a href="Example of Answers to Obtain Tweeter Developer.pdf"> See Examples of the Answers for the Questions from Twitter to Get a Twitter Developer's Account //a href="SampleAnswersForStudent_TwitterDeveloperAccountApplication.pdf"> (This is For an Academic Research Account, which You Don't Need to Apply) See Sample Answers for the New Questions from Twitter to Apply Twitter Developer's Account For Your Twitter Account Application For the Project Site if Asked, Provide Your Lab3 Specification above. you Can also Provide the CIS612 Project Site and Research Project Description for Big Data and Data Scientist //a href="https://eecs.csuohio.edu/~sschung/cis612/CIS612Project.html"> CIS 612 Project Site to Provide in Your Application for Twitter Developer's Account //a href="Faculty_Led_Project_Description__Social MeadiaSentimentAnalysisSystem_SunnieChung.pdf"> To Answer with Sample Research Project Description for Big Data and Data Scientist How to Get Twitter Stream in JSON file: //a href="CollectingTweetStreamingDataintoJSONfileWithPython_Sp2022.pdf"> Step by Step Guide on How to Get Twitter Stream in JSON file in Python (After the Tweepy API Version 4.0 as of Spring 2022) NEW POST !! by TA Yixi Luo **************** //a href="Fetching tweets data from Twitter with Python.pdf"> Step by Step Guide on How to Get Twitter Stream in JSON file in Python (Before The Tweepy New Version 4.0) How to Collect a Twitter Stream to a JSON File then Insert to MongoDB //a href="StepByStepInstructions_CollectingTwitterStreamFromDeveloperAccount.pdf"> Tutorial on How to Get Twitter Stream to Insert into MongoDB in Python As of 2021 Before the New Tweepy Version 4.0 Note: This tutorial uses last year (2021)'s Tweepy Streaming Library. It has been upgraded to a new version as of early 2022. See The Tutorial for the New Version of Tweepy Streaming Lib Above Blindly copy and paste of the codes in this tutorial wouldn't work because of the new version of Tweepy Streaming Lib //a href="https://eecs.csuohio.edu/~sschung/CIS593/CIS593_Lab2_JSON_FAQs.pdf"> FAQs for JSON Processing in General //a href="FetchingTweetsfromTwitteronHDFSMongoDB.pdf"> Other Related Documentations and Examples from Twitter sites and Python for Twitter Visualization Tools of GEO Spatial Data -- NOT Required for Everyone !! //a href="https://www.qgis.org/en/site/"> QGIS //a href="https://docs.qgis.org/3.10/en/docs/pyqgis_developer_cookbook/intro.html#scripting-in-the-python-console"> QGIS in Python //a href="https://eecs.csuohio.edu/~sschung/cis612/How%20to%20Visualize%20Geospatial%20Data%20Using%20MAP%20API%20QGIS.pdf"> Example to Visualize in QGIS //a href="https://community.esri.com/thread/202728-why-are-all-grids-displaying-in-the-ocean-near-africa"> Errors: Why Data Points are Mapped into Ocean? //a href="NIJClusteringHotSpotAnalysisusingPython.pdf"> Example of Lab2 GIS Data Processing in Python: Example Project to Guide How to Process Geo Spatial Data (generated from a machine) for Hot Spot Analysis From CIS 660 Project by Sarvesh Chande //a href="ArcGIS_NIJData_Dan.pdf"> Lab2 GIS Data Processing Example of NIJ Geo Spatial Data Visulaization Using ArcGIS Map for Hot Spot Analysis NEW POST ! //a href="CIS 612Lab2ExtraNIJGeoSpatialArcGIS_Sriram.pdf"> Example of NIJ Geo Spatial Data Visulalization Using ArcGIS 3D Map for Hot Spot Analysis NEW POST ! //a href="NIJGeodataVisulaization_JavaScriptHTML_Heta.pdf"> Example of GIS Data Processing in Java Script to Visualize in HTML and //a href="NIJGeojsonDataFile.js"> NIJ Data file in GeoJSON NEW POST ! //a href="https://www.expertgps.com/spcs/Oregon-North-FIPS-3601-NAD83.asp"> Example of Coordinate Conversion to Oregon North FIPS 3601 NAD83 NEW POST ! //a href="NIJGeoDataClusteringReza.pdf"> Example of NIJ GIS Data Processing in a Plain Coordinate in Python for Clustering Analysis //a href="ClusterNIJGeodataHeta.pdf"> Example of NIJ GIS Data Processing in a Plain Coordinate in R for Clustering Analysis Other Data Set -- This is NOT for This Semester //a href="CIS612Lab3_3JSONYelpDataWithExtraBusinessDataTransformation.pdf"> Lab3_3 on GEOJSON Processing: Use GEOJSON data file below instead of Yelp Business Data //a href="https://eecs.csuohio.edu/~sschung/CIS660/Cuyahoga.zip"> GEOJSON Data File for Lab3_3 on GEOJSON Processing (~ 1GB Zip file) //a href="GEOJSON Structure.pdf"> GEOJSON Data Description See the Class Lecture Notes Section for MongoDB Guides and Set up, and More Details to Learn MongoDB CRUD Operations Lab 4: You can choose either Lab4_1 or Lab4_2 below 1. Building an Simplified Inverted Index in a SQL Server for Lab4_1 (Minimum Required) For Inverted Index to Build, You Can Simplify to One Table with (Term, Doc#, TermFreq) 2. Building a Full Inverted Index Either in a SQL Server or MongoDB as in the Lecture Note as below: (Extra Credit) Dictionary table (Term, TotalDocsFreq, TotalCollectionFreq) and Posting Table(Term, Doc#, Term_Freq) 3. Extra Credit: Build Inverted Index with NLP Pipelining for each Sentence to Extract Context Aware Information with POS or/and NER Tagger and Store them either in SQL Server or MongoDB for Retrieval Later Input Files for Lab4_1 //a href="https://eecs.csuohio.edu/~sschung/CIS593/Text_Mining_ConstructingInvertedIndex.pdf"> Example of Inverted Index on State Union Addresses (Note that the Structures of the Index Tables are a little Different in the Example) See Extra Credit Lab 4_2 on Information Extraction of Biomedical Document Collection using STANZA NER and OpenIE to Extract TRIPLES to Create a Knowledgebase in JSON/MongoDB Collection Big Data Set and Tools: //a href="https://dumps.wikimedia.org/"> wiki Data set in XML and Html //a href="https://meta.wikimedia.org/wiki/Datasets"> Wiki Text Data Sets Complete List and Tools //a href="https://pubmed.ncbi.nlm.nih.gov/download/#annual-baseline"> Pubmed site XML Description //a href="https://ftp.ncbi.nlm.nih.gov/pubmed/baseline/"> 32 million Paper Abstracts in XML in Pubmed Baseline Text Preprocessing Library in Python SpaCy: //a href="https://spacy.io/api/lemmatizer"> Lemmatizer in Python SpaCy //a href="https://stackabuse.com/python-for-nlp-tokenization-stemming-and-lemmatization-with-spacy-library/"> Liquistic Modules for Tokenization, Stemming, Lemmatization in Python SpaCy //a href="https://stackoverflow.com/questions/38763007/how-to-use-spacy-lemmatizer-to-get-a-word-into-basic-form"> How to Code Liquistic Modules like Lemmatizer in Python SpaCy //a href="https://codeburst.io/python-basics-11-word-count-filter-out-punctuation-dictionary-manipulation-and-sorting-lists-3f6c55420855"> Python Example for Basic Text Processing (NOT For This Semester !!!!!) Lab5 Sentiment Analysis with Machine Learning for Classification: Data Set: Choose a Review Text Data Set Obtained from Social Network sites: Twitter or Yelp Review Data Set Classification Goal: 1. For each review text obtained from Twitter texts, Derive a Preditive Model to Predict (Classify) Whether it is Positive or Negative Review. 2. For each review text in Yelp Review Data Set, Derive a Preditive Model to Predict (Classify) the scale of the Review in 1 - 5. (star) Useful Big Data Analytic Tools Choose your System/Tool/Platform to Set Up and Get Used to: Python Analytics Tools and Tutorials : Machine Learning : Other Machine Learning Platforms: Basic R Tutorials : Lab 5: Data For Hive or Pig Latin: Use either Video Game Sales Data below or Yelp Business.json For Those who Want to Set Up Your HDFS Cluster On EC2 Amazon Cloud, See the Cloud Section at the end of the Class Lecture Note Section for a Student Account. Hadoop Set Up Instruction Sites: //a href="http://hadoop.apache.org/docs/r1.0.4/single_node_setup.html"> Hadoop single node setup //a href="http://hadoop.apache.org/docs/r1.0.4/cluster_setup.html"> Hadoop Cluster setup //a href="http://hadoop.apache.org/docs/r1.0.4/mapred_tutorial.html"> Map Reduce Tutorial on Hadoop Data Set for MR Job: //a href="access_log_Jul95.zip">NASA HTTP Access Log File //a href="http://icsdweb.aegean.gr/awid/features.html">AWID:Wireless Network Server Log Data Set AWID Data Set Avaliable here as well: //a href="https://drive.google.com/open?id=0ByArNbFEPXXxZGo1bDJfSGRpbms">AWID: Wireless Network Server Log Data (around 20 GB zip) //a href="https://aws.amazon.com/docker/"> You can use a Container: Docker (Instead of VM) to Configure a Distributed Cluster for Lab4_1 Lab Guides: Newest on the Top //a href="Instruction_INSTALLING_HADOOP_Ubuntu.pdf"> Hadoop Installation on Ubuntu: How to Install Hadoop on Ubuntu (as of 2020) //a href="TroubleShootingTipswithHadoop.pdf"> Trouble Shooting Tips for Installing Hadoop on VM Do not reformat again to avoid losing name node data //a href="Lab4_1InstallHadoopDataNodeError.pdf"> Hadoop Installation: How to Fix When Data Nodes are not Running (2018) //a href="https://stackoverflow.com/questions/11889261/datanode-process-not-running-in-hadoop"> Help Site on How to Fix When Data Node are not Running (2018) //a href="HowtoExecuteMapReduceinEclipse.pdf"> Procedure to How to Execute MapReduce in Eclipse to Run a Wordcount Job (2018) //a href="Lab_4_1_Alex_Chengelis.pdf"> Procedure to Install Hadoop and Run a Wordcount Job on Window 2010 (2017) For Set up Problems, The following post is helpful. //a href="http://stackoverflow.com/questions/21005643/container-is-running-beyond-memory-limits"> Container is running beyond memory limits A good Instruction Site for Installing and Running Hadoop //a href="Setting-up-Hadoop-made-easy.pdf"> Installing Hadoop If you have a trouble installing Hadoop from the above site with not seeing the name node, You need to delete the temp files created in standalone mode and reformat the namenode. For setting up a passwordless SSH, see Ganesh's Lab4 below as well. //a href="HadoopAssignment4Ganesh.pdf"> Installing Hadoop and running a wordcount job by Ganesh VAVILAPALLI //a href="https://www.youtube.com/watch?v=MoKW5eY5yVY"> Video for Installing Hadoop shared by Prashant Patel //a href="Lab_MRHadoop_Um.pdf"> Guideline3 for Lab4_1 on Mac (2013) MR Programming on Hadoop Lab: (NOT FOR THIS SEMESTER !!!!) //a href="CIS 612_LabAssignment4_2_InstallNoSql.pdf"> Lab4_2 on MR Programming on Hadoop: Implement Average Temp By Station in MR on Hadoop as in Lab Guide below. Follow the sample codes below. //a href="MapReduceHadoopLabs.pdf"> Simple MR Hadoop Lab Guides //a href="MRsampleCodes_AverageTempByZipCode.pdf"> MR sample Codes - Average Temp By Zip Code //a href="MRSampleCodes_SortByStationID.pdf"> MR Sample Codes - Sort By Station ID //a href="stationData_5000.txt"> Zip file for station Data Set and //a href="Lab5MR.zip"> Zip file for station Data Set and MR Codes NoSQL Systems on Hadoop Lab: Part 1: Due By the midnight of the End of the Fourth Weekend of November Part 2: Due By the midnight of the End of the First Weekend of December All the Extra Credit Labs By the end of the Last week of the class in December! You can choose One Parallel NOSQL System for Lab5_2 below. //a href="CIS612_Lab_4_2_NoSQLHivePigSpark.pdf"> Lab5_2 on Using NoSQL Systems: Hive, PIG, MongoDB, HBase, or Spark For Part 2: Data For Hive or Pig Latin: Use either Video Game Sales Data below or Yelp Business.json //a href="http://myitlearnings.com/creating-hive-table-partitioned-by-multiple-columns-and-importing-data/"> Example of Creating a Hive Table with Partitioning //a href="https://www.guru99.com/hive-partitions-buckets-example.html"> Required Setting for Partition and Step by Step Example for Creating Hive Partitioning Tables with Partition By Note that For Creation of Partition Tables, for partition, you have to set this property set hive.exec.dynamic.partition.mode=nonstrict Note that Hive has a bug that initialize your Name node whenever you start Hive again. You may lose your data so BACK UP Your Data ! Most of set up problems are coming from a Version Mismatch between conponents. It may not have documented correctly (remember these are open sources). Distributed Parallel NoSQL System Set Up and Basic CRUD Operations in Examples : //a href="CIS612_Lab_4_2_Tutorials_HiveTablePatitionsForDataWarehouse.pdf"> Hive Table Patitions For Data Warehouse //a href="CIS612_Lab_4_2_Part2_MongoDBJoinScript_Asanka.pdf"> MongoDB Join in Python Script as Client //a href="CIS612_SparkBasicProcessingTutorialHeideloff.pdf"> Spark Basic Data Processing //a href="CIS612-LAB4_2_HiveHBaseJoin_Paul.pdf"> HBase Join with Hive Hive Set Up Guides for Lab4_2 (2019-2020) //a href="CIS612_Lab4_2_Hive_Installation.pdf"> Hive Set Up guide (2020) NEW POST !! You need to use new version JDK for a newer HIVE version: oracle – 8 -jdk (istead of using default JDK) //a href="CIS612_Lab4_2_Hive_CommonInstallationProblems.pdf"> Common Hive Set Up Probelms and Solutions NEW POST !! //a href="HiveInstallationGuide.pdf"> Hive Installation Guide (2016) Data SetS: //a href="videogamdata.csv"> Video Game sales Data for HIVE/PIG If You choose MongoDB, Do Aggregate Data Pipeling in Lab3_4 in the Lab3_4 Section Above. Data for Mongo DB: Use business.json and review.json from the Yelp site //a href="https://www.yelp.com/dataset"> JSON Files from Yelp Challenge //a href="access_log_Jul95.zip">NASA HTTP Access Log File //a href="http://icsdweb.aegean.gr/awid/features.html">AWID:Wireless Network Server Log Data Set AWID Data Set Avaliable here as well: //a href="https://drive.google.com/open?id=0ByArNbFEPXXxZGo1bDJfSGRpbms">AWID: Wireless Network Server Log Data (around 20 GB zip) NOSQL System Set Up Guides for Lab4_2 //a href="CIS612_Lab4_2_Hive_Installation.pdf"> Hive Set Up guide (2020) NEW POST !! You need to use new version JDK for a newer HIVE version: oracle – 8 -jdk (istead of using default JDK) //a href="HIVE-HBASE-INSTALL-TUTORIAL.pdf"> HBase with Hive Set Up guide //a href="Hadoop_Cloudera_Hive_SetUp.pdf"> Hadoop/Cloudera/Hive Installation Guide NOSQL System Set Up Guides for Lab5_2 //a href="CIS612LAB4_2_MongoHiveKeshav.pdf"> MongDB/Hive Installation Guide //a href="Lab4_2_HiveMongoDBSonal.pdf"> MongoDB/Hive Installation Guide, Permission Error Resolution //a href="Lab4_2HiveMongDB_Sagar.pdf"> MongoDB/Hive Installation Guide //a href="Lab4_2_PIGMongoDBSamantha.pdf"> PIG/MongoDB Set Up Guide Spark: Real Time System Set Up Guides: //a href="CIS612_SparkBasicProcessingTutorialHeideloff.pdf"> Spark Installation Guide with HDFS (When HDFS is already set up on your system) //a href="http://spark.apache.org/docs/latest/hadoop-provided.html"> Spark Installation/Configuration Guide with HDFS //a href="CIS612_SparkSetUp_HDFS.pdf"> Spark Set Up Guide with HDFS //a href="CIS612_SparkInstallation_Ubuntu.pdf"> Spark Installation Guide on Ubuntu (Without HDFS) //a href="Instruction_Installation_Spark_Win10.pdf"> Spark Installation Guide on Win10 Kafka: Real Time Messaging System Set Up Guides: (2020) //a href="CIS612_Kafka_Installation_Ubuntu.pdf"> Kafka Installation Guide on Ubuntu Cloudrea: Integrated Real Time Data Analytic Platform: SparkSQL, Spark, Hive on Hadoop //a href="Install_Cloudera2018.pdf"> Cloudera Installation Guide NoSql System Set Up Guides for Lab5_2: //a href="HBaseCassandraInstallationGuide.pdf"> HBase and Cassandra Installation Guide //a href="HBaseTutorial.pdf"> HBase on Cloudera Setting up and Tutorial //a href="SparkTutorial.pdf"> Spark Tutorial From Danielle Aring's Report //a href="VoltDBCommandSPSampleCodes.pdf"> VoltDB (In-Memory Relational Database Server) Commands and Stored Procedure Sample Codes You can create your Big Data Processing Infrastructure on Amazon Cloud for your Project //a href="http://aws.amazon.com/education/awseducate/"> Amazon Cloud Account for Students Supporting Contents for Labs : How to Create a Web Application with Java Based Application Server with MS SQL Server: Amazon Cloud: //a href="https://aws.amazon.com/rds/">Amazon RDS (Relational Database Service) //a href="http://aws.amazon.com/ec2/">Amazon Elastic Cloud Computing (EC2) for Web service //a href="http://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/TUT_WebAppWithRDS.html">How to Create Amazon Virtual Hosting (EC2) with a Web Server and Amazon Database Server (RDS) Microsoft Cloud Azure: //a href="https://azure.microsoft.com/en-us/">Trial account for Microsoft AZURE Cloud //a href="MicrosoftAzureTables.pdf">How to Create/Retrieve a Table in Microsoft AZURE Cloud //a href="https://code.msdn.microsoft.com/Fix-It-app-for-Building-cdd80df4"> Sample Project on Microsoft AZURE Cloud //a href="http://www.asp.net/aspnet/overview/developing-apps-with-windows-azure/building-real-world-cloud-apps-with-windows-azure/unstructured-blob-storage">How to Create/Retrieve BLOB data in Microsoft AZURE Cloud //a href="http://www.davidchappell.com/Azure_Services_Platform_v1.1--Chappell.pdf">AZURE Cloud //a href="windows_azure_sql_database_tutorials.pdf">How to Create a SQL Database Server in Microsoft AZURE Cloud //a href="http://www.tutorialspoint.com/microsoft_azure/index.htm">Tutorial for MS AZURE Cloud Useful Resource Sites: |
| Class | Chapter / Topic / Specific Objectives / Activities |
| 1 |
|
| 2-5 |
|
| 7-8 |
|
| 10-11 |
|
| 5-6 |
|
| 12-13 |
|
| 13-14 |
|
| 13-14 |
|
| 15 |
|
| 15 |
|
| 15 |
|
| 6-7 |
|
| 16 |
//a href="ENACh13final-Disks-FileStructure-Hashing.pdf"> Project Presentation |
==> Completion of Homeworks/Labs is required for obtaining a passing grade.
| This
is a tentative scale and |
Letter |
Quality Points |
|
||
| A |
> 93% |
A: Outstanding (student's performance is genuinely excellent) | |||
| A- |
90% - 93% |
||||
| B+ |
87% - 90% |
||||
| B |
82% - 87% |
B: Very Good (student's performance is clearly commendable but not necessarily outstanding) | |||
|
|
B- |
80% - 82% |
|||
|
|
C |
75% - 80% |
C: Good (student's performance meets every course requirement and is acceptable; not distinguished) | ||
| D | 65%-75% | D: Below Average (student's performance fails to meet course objectives and standards) | |||
|
|
F |
<65% |
F: Failure (student's performance is unacceptable) | ||
|
ADA Adherence. If you need course adaptations or accommodations because of a disability, if you have emergency medical information to share with me, or if you need special arrangements in case the building must be evacuated, please make an appointment with me as soon as possible. My office location and hours are listed on top of this syllabus. If you need further information, please contact the ACCESS office, phone number 687-5106. |
Programming standards
|
|
|