[0x01] – Introduction and Overview

Welcome hackerz,

It’s been a while since my last post. I’ve been busy dealing with lots of unforeseen problems IRL, and additionally, I’ve dedicated the time I did have online towards finishing off an exploit chain I had been development (I wanted to sell it before someone else beat me to it). Now that I’m back with this new series, you can begin to expect regular posts again.

Note that this is going to be a very long tutorial series – fifteen separate blog posts, in fact. Don’t be discouraged if you haven’t learned much after the first few posts. To benefit fully from this series, you’ll have to pace yourself by reading through all 15 posts, and additionally you’ll need to play around with the downloadable tools I listed below (I will have some challenges involving those tools in upcoming instalments of this tutorial series.

For this introductory post, I am just going to be explaining what the series is about, explaining some concepts that are unique to C that you’ll need to familiarize yourself with, explaining different data types within C, and finally showing you what the basic structure of a C program looks like, and the purpose of each part of that structure. In Part #2 I will be moving on to stuff like writing actual syntax, usage of particular functions, uses of conditional logic, how to safely perform dynamic memory allocation, advanced compilation tips and tricks, and so on. If you plan to follow this series the whole way throughout, then please look into the tools and resources I suggested, as some of these will be needed to complete challenges that I will be setting throughout the series. The idea of this series is that if you read all 15 posts, while also completing the challenges that I set, then by the end of it all, you should be somewhat fluent at writing code in C/C++ and Assembly.

This series will be split into five primary sections:

  • C/C++ programming introduction and fundamental concepts (Posts #1, #2, and #3)
  • C/C++ socket programming, and intermediate/advanced concepts (Posts #4, #5, and #6)
  • Explanation of computer architecture and CPU fundamentals (Posts #7, #8, and #9)
  • RTL (Register Transfer Lang) tutorial, and introduction to JASP Toolkit (Posts #10, #11, and #12)
  • Tutorial on 16bit ASM via JASP, introduction/fumdanentals + tutorial on x86 (Posts #13, #14, and #15)

For an overview of what will be covered in this series across all five sections, see the list below:

this list displays a number of different key concepts that will be discussed throughout this 15-part tutorial series (in no particular order)

  • which compiler and/or IDE to use
  • The basic fundamental concepts of C/C++
  • C/C++ code examples
  • C/C++ Linux vs Windows
  • Introduction to memory (stack, buffer, etc)
  • Guide to proper memory allocation
  • Introduction to bitwise operators
  • arithmetic quirks
  • more advanced concepts within C/C++
  • socket programming within C/C++
  • advanced compilation options (different flags etc)
  • Some more intermediate/advanced code examples of C/C++
  • Introduction to ASM basics
  • List of common opcodes and their purposes
  • basic explanation of CPU architecture, endianess, etc.
  • Introduction to RTL (Register Transfer Language)
  • Installing and setting up JASPer (16bit CPU emulator)
  • Learning how to use JASP, passing it opcodes for 16bit ASM, etc
  • Back to RTL, but more advanced/complex pseudo-code snippets this time
  • Basics of x86 Assembly
  • writing in-line Assembly for C/C++
  • Intermediate x86 assembly
  • suggestions as to what you should cover next
  • final notes

At first (especially if you have no prior knowledge in the field), this is DEFINITELY going to feel overwhelming. This isn’t something you can become competent at over the course of a weekend or a week. That being said, once you do become competent at this stuff, you’ll be astonished at how incredibly powerful it can be.

Even if you don’t have programming knowledge whatsoever, I still think it could be very beneficial and insightful for you to read this guide, and follow along with it as best as you can (while also taking part in the interactive testing exercises I’ll be setting up via JASP toolkit throughout several parts of this journey) . I’m sure scores of programmers would strongly disagree with mw in saying that you should still attempt to read this guide even if you have no prior pogroming knowledge. People instead suggest that you start with an “easy” language that is interpreted rather than compiled, such as python for example… while this isn’t necessarily bad advice I’m still going to go with my own route, and make my first programming tutorial series cover the langs I learned by diving right into the deep end. People say languages such as python are a good choice because they’re easy, which, sure, they are… but, they rely on lots of premade third-party libs for functions, rather than you having to manually write your own function then call it when necessary, like you would in C… or, you can assign whatever you want to memory in any manner you want while you’re writing python code, but try that in C, and someone is 100% going to spot the security flaws in your code within seconds. I think by starting with C, it teaches sensible and clean syntax, it teaches how to think like a programmer a lot more than many interpreted programming languages do. It teaches students the fundamental concepts of CPU architecture, alongside other useful concepts such as dynamic memory allocation or manual construction of sockets for client/server communications. Sure, you can automate much of this in python (or even ignore it entirely when it’s regarding issues affecting memory management), but by learning how to do it manually, you’re going to gain a much deeper understanding of how these concepts work on a fundamental level. It is also worth remembering that a hugely vast number of languages follow a C-style syntax and similar rules to C, so if you’re intending to learn more than one programming language, while starting out with something like C is going to be challenging, the long-term payoff makes it worthwhile, as it allows you to pick up other languages in future MUCH EASIER than you’d have picked them up without knowing how to write and read code in the underlying language that is the basis for many of these modern languages.

This post will be the first post covering part of section one — I will be offering a basic introduction to the C programming language, followed by explanations of many of the fundamental core concepts within C, accompanied by code examples to demonstrate each of these contexts.

Oh, also… despite the fact I am going to be teaching C++ within this series too, please for the love of god try to keep your syntax native to C wherever possible (although compile your C code via g++ for the reasons I list within the next chapter). C++ will encourage bad programming habits, and itt’l cause you to rely too heavily on third-party libs and functions, while having no idea how you’d go about manually implementing such functions within regular C. If you don’t want to take my word on this, then listen to the god of Linux himself. Or take a look at this wonderful list of quotes from industry professionals.

Maybe they’re bashing on C++ too hard, it definitely does have its benefits. For example I think C++ would be a more logical choice than C for game development.. but, for the most part, C is just inherently better than C++. Anything that can be done in C++ can be done in a more efficient manner, using regular C.

Anyhow, lets dive right into things!

[0x02] – What you need to know before you begin:

First, I will give a quick list of every tool and resouce you’re going to need for this entire series (all 15 posts). This way, if you’re missing anything off the list, you can download and install it prior to me releasing the next part of the series).

The following tools and resources will be required throughout this series:

  • A Windows machine or Windows VM (win7 or above).
  • A Linux-based machine or VM (Any modern distro will do just fine).
  • gdb, ollydbg, ghidra, or some other form of debugger
  • JASP Toolkit (http://www.brittunculi.com/jasp/)
  • Sedici Toolkit (http://www.brittunculi.com/sedici/)
  • An IDE (CodeBlocks for example) or at least a decent text-editor (Sublime or Notepad++ for example)
  • gcc AND g++ ready to be ran within your Linux environment
  • valgrind
  • git (you don’t need a github account, just the ability to git –clone and so on)

The following resources are not a necessity, but they are definitely useful and will speed up your learning process along the way, while encouraging good coding habits:

Before you even consider starting, you need to come up with some sort of environment for you to be writing your code in – for example you should pick which compiler you decide to use (more on this later), and additionally you should find an IDE that you enjoy using, or at the very least find a decent text-editor.

Personally, I’ll use nano or vim if I’m writing my code within an SSH session into one of my boxes, but if I’m writing my C/C++ code on a desktop environment, then I’ll use a text editor that actually has some half-decent features. I’ll often stick with Sublime or Notepad++, but of course this is all just a matter of personal preference. Find what best suits you, and stick with that for the time being. It doesn’t matter too much what software you’re using to write your code — what is really important is that you’re actually spending the time and effort to write code!

As for your compiler, there are tons of options available – once again, it’s a matter of personal preference. As for myself, I use MinGW if I’m on Windows, which is essentially as close you’re going to get to GCC/G++ without being on a Linux system.. and, speaking of Linux systems, I will use either GCC or G++ for code compilation on Linux.

Something worth noting, is that even if you are writing your code in C, it is best to compile as C++ (so, via G++ rather than GCC — while it is still possible to compile C++ files via GCC, they’d need to have a .cpp extension, whereas G++ on the other hand will compile standard C files with a .c extension as C++), there are a handful reasons for this:

  1. G++ has faster compilation speeds compared to GCC
  2. G++ has better debugging and error handling capabilities
  3. G++ not only has better debugging and error handling capabilities, but it also provides more extensive error message outputs compared to GCC, allowing you to diagnose the issue with more ease.
  4. By compiling in C++, you are able to take advantage of the bool data type. Regular C on the other hand does not support boolean data types, therefore to access that data type you’d need to compile it as C++ rather than C.
  5. G++ allows linking to object files, whereas GCC has no support for this
  6. G++ comes with far more defined macros

In post #2 of this 15-part series, I will be discussing compilation in more depth. Particularly I’ll be discussing the purposes of specific flags that can be invoked while calling GCC or G++, and the purposes of each of those flags, alongside a few handy tricks to speed up compilation times, among other things.


[0x03] – What makes C/C++ differ from other programming languages?

Well, where do I even begin? If I listed everything that makes this different from say, “Java”, for example.. then this 15-part series would turn into a 25-part series in almost no time at all.

The very first thing to remember about C (this may confuse some python programmers) is that spacing and indentation make no difference as to how your code is wrong. They have no relation whatsoever to the syntax. As long as the code you write is syntactically correct, you could have a fully-working C program with thousands of variables and hundreds of custom functions, all within a single line in your text edit — it would still compile without any issues, because spaces don’t mean a thing here. Contrast that with Python, where spaces and indentation are used to dictate conditional logic and actually control the flow of the program.

For example, if you were to do an if statement in Python, it would be done like so:

#!/usr/bin/env python
nolifer = True
if nolifer==True:
    print("Welcome to my basement.")
    print("Care for any cheetos or tendies?")

As you can see, no curly braces were harmed during the making of this code. Python relies on spacing and identation to dictate conditional logic, so rather than wrapping curly-braces around the code they want to trigger if the if statement condition is met, they use either the TAB char (or the SPACE char four times) to intend the snippet of code that they want to execute if the condition is met.

Compare this to C or C++ which relies entirely on curly-braces, where spacing/indentation doesn’t mean a thing:

#include <stdio.h>
int main() {
  int leet = 1337
    if (leet == 1337) {
        printf("DEM THICC CURLY BOIS");
    }
    else {
	printf("Programmers against identation! CURLY BOIZ 4 LYF!");
    }
}

So, if you’re used to Python and you’re reading this tutorial, that’s one thing you’ll notice almost immediately. Spaces don’t mean anything in C, it’s all about using curly-braces to dictate your conditional logic instead. Due to spaces literally not meaning a thing, you can just outright ignore them entirely:

#include <stdio.h> int main(){int one=1; if(1=1){printf("one is one\n");}else{printf("lol");}return 0;}

The above code works exactly like it would if it was properly formatted. Please don’t misconstrue what I’m writing here, by the way. You absolutely should still be using spaces and indentation while writing your C code – despite them not having any practical or functional use, nor any real meaning to the compiler… they should definitely still always be included, because your code will always look a lot cleaner that way, and itt’l be far far easier to read – not to mention itt’l get you into the habit of writing clean code over a long term period, and honestly, writing clean code can be just as important as writing efficient code.

One interesting feature about C/C++ is these are multi-paradigm lanauges, rather than being limited to a single paradigm as is the case with many other languages. C takes influence from languages that follow the imperative or procedural styles of programming paradigm, whereas C++ takes influence from those also, but is a lot more leaning towards event-driven and object-orientation types of paradigm. The fact that these languages support a bunch of different paradigm means that a paradigm of choice can be implemented based on the unique style of whoever is writing some code – this is yet another reason why I think C/C++ are great languages to learn. If you’re learning a language that relies on a single paradigm, then your coding style is somewhat limited… but if you’re learning a language that easily allows for multiple paradigms to be implemented, it gives you a better understanding of the pros and cons to each type of paradigm, and allows you to become well-versed with writing code in a variety of different manners.

Two major differences in C/C++ that you aren’t going to have to deal with in other languages, are pointers, and dynamic memory allocation. A pointer is simply a special type of variable (although all variables can be treated as pointers) that “points” to the memory address of a variable. For example, when you’re accessing a variable, you will be seeing the data that is stored in memory for that variable (for example the value that was assigned to the variable), but if you access a pointer to that variable, you are instead seeing the memory address of the variable itself, as opposed to the data stored within the variable. To set a pointer, you just use the * prefix, preceding the variable name and after the data type declaration. For example, accessing int one would access the value stored in the variable “one”, but accessing int *one would give you the memory address for where the “one” variable is stored.

Here is an example of how a pointer would be declared, along with some comments to explain what is going on (this will be covered far more in-depth in part two):

int *ptr;     // declare integer pointer
int num = 4;  // declare variable that holds the value 4
ptr = #   // Assign the address of the "number" variable to the pointer
*ptr = 5;     // Assign the value 5 to the address contained within "pointer",
              // which at the moment is the address of "number", which changes the value of "number" to 5 

While on the subject of memory, it’s worth noting that C is a “memory unsafe” language, meaning that you need to manually deal with dynamic memory allocation otherwise it will result in issues such as buffer overflows. I’ll be going in-depth in regards to different memory types, and safe methods of proper memory allocation within parts 2 and 3 of this series.

Another big difference one may notice is how variables are dealt with in C/C++. Depending on what languages you have worked on in the past, you may have either come across “strongly typed” languages, or “weakly typed” languages. Now, there are a few specific caveats that make C very hard to categorize as either strongly or weakly-typed. Before I get into that though, I’ll give an explanation for example. Within PHP, if you are specifying a variable, you can just specify it like so:

<?php 
  $input = $_GET['user_input'];
    echo "The value of the user input is". $input. " - No we will check the data type:";
    echo "<br />"
    echo "data type:".gettype($input);
?>

So, for those who don’t know PHP, all the above script is doing is taking a user-supplied URL input, like so: http://example.com/enter-input?user_input=here – the user-supplied input is the assigned to the $input variable value. I am then outputting the value of the gettype(); function with my variable as an argument. This will output to the screen that my input was of the “string” data type, since I applied the string “here” as the user input… but, if I were to instead supply 1337 as the user input, then the output of this PHP script would be “integer” rather than string. Variables here don’t need to have a data type pre-specified… instead, they automatically determine the correct data type based upon the input passed to them. This makes PHP a “weakly-typed” (also known as loosely-typed) language. Another definition for strongly vs weakly typed is “statically” vs “dynamically” typed – statically-typed languages will perform type checking at compilation type, and dynamically-typed languages will perform those checks at runtime. Although these terms are often used interchangeably, although there is in fact a difference between them. For example, a dynamically-typed language can also be strongly-typed… people say that a language being statically-typed (i.e. type checks at compilation) automatically makes it strongly-typed, but that is most certainly not the case, I’m unsure where this misconception comes from, but it likely stems from the lack of standardized definitions. As you’ll see in the example below, C is statically-typed, and you’d also think it was strongly-typed based on the fact that you have to specify a data type when declaring or assigning a variable value, but, as you’re about to see, it’s actually more complicated than that.

as stated, within C, you do have to specify the data type in use, like so:

#include <stdio.h>

int one = 1;
char M = "M";
float leet = 1.3.3.7;

int main() {
printf("%d\n", one);
printf("%c\n", M);
printf("%d\n", leet);
}

C is a tricky one… many would argue that it is “strongly-typed”, myself included for the most part… but, there is a fair amount of room for interpretation in the definitions of “strongly-typed” vs “weakly-typed” and many people would argue that C is technically a weakly-typed language since it explicitly permits the typecasting of pointer values. Personally, I think the best way to categorize it is by thinking of this as a spectrum, rather than in terms of black and white. I’d consider it a strongly-typed language, but LESS strongly-typed than other languages. I wouldn’t go as far to call it weakly-typed. Maybe just “not-so-strongly-typed“.

The following data types are available within C:

  • char (used to store a single character)
  • int (used to store a whole number)
  • float (used to store floating point decimal numbers with “single” precision)
  • double (used to store floating point decimal numbers with “double” precision)

a data type involving numbers can be both signed or unsigned – the former means its variable value can support negative numbers, whereas the latter means it cannot.

There are also different ranges for minimum and maximum size for values, for example you can use short int for an integer value with a shorter range, or, adversely, you can use long int or even long long int to allow your integer value to have a wider accepted range.

You may have noticed that there is no “string” data type – that’s because technically, there’s no such thing as strings in C — Instead, you use the “char” data type to create a one-dimensional array of characters spelling out the string you want (you need to ensure that the array size is one byte larger than the size of the string itself, this is because you need an extra byte to account for the “null-terminator char” which acts as an escape sequence and essentially tells the compiler that the string has ended). an example of the string “0xffff” could be demonstrated like so (with the \0 at the end specifying the null-terminator:

char 0xffff[7] = {"0", "x", "f", "f", "f", "f", "\0"};

There are a number of other ways to represent strings within C, but, as stated earlier, the purpose of this post is only to introduce some very basic fundamentals. I’m not going to be getting into the nuances of specific syntax or giving actual source code examples until the second instalment of this 15-part blog series (for anyone confused as to what i meant by “one-dimensional array”, I will also be explaining array types in far more detail in Part 2).

Remember, once again, C is a compiled language not an interpreted one. With a scripting language such as Python, the code is fed to an interpreter which will interpret (hence the name) the code in real-time, whereas C/C++ will be built into an executable by the compiler. The compilation process first converts the C into Assembly language, and from there it is converted directly into machine code, allowing it to be executed correctly by the machine.

Now that I’ve discussed some basics of the paradigm (something I’m going to dive into properly in part 2 of 15 — along with a much more in-depth explanation of the compilation process, including the uses for different flags in gcc/g++), we will move onto the basic structure of what a C program looks like.

[0x04] – How is a C program structured?

Before moving onto part two where we will begin covering basic syntax, conditional/logical operators, bitwise functions, memory allocation, and so on.. I am first going to finish part one by giving a general overview of how a C program is structured/formatted, and the purpose of each section:

Thanks for reading so far! Now that I’ve got some of the bare basics covered, I am going to start releasing actual code examples and teaching syntax + conditional logic and so on, for part 2 of this series. In Part 3, I will be moving onto socket programming.

Apologies if Part 1 wasn’t very technical at all.. I’m trying to keep it as noob-friendly as possible to ease people into it slowly, since we’ll be diving right into the deep end by the end of Part 2 and start of Part 3.

That’s all for now. Part 2 with actual syntax tutorials + working code examples will be coming tomorrow.

./mlt –out