Login Register






How hackers exploit flaws in code ? filter_list
Author
Message
How hackers exploit flaws in code ? #1
Question asked on reddit:
Quote:What is exactly does a coder do to exploit a flaw in code? This may be out of the scope of programming, but I ask because I read that phones are unlocked by gaining a serial connection to a piece of hardware, then finding flaws in the kernel that allow them to gain root access.
But how does one exploit? Are they able to upload their own code somehow?
To bring it back to more to programming, how do the previous questions apply to hackers that crack software. Do they have to obtain the entire source code for a piece of software to somehow modify the code and recompile it? If so, how did they get the code in the first place?
I'm very curious about this. Please help me understand this.


Heartbleed is more like an info leak. SQL injection and XSS have nothing to do with rooting phones either.

Quote:But how does one exploit? Are they able to upload their own code somehow?

Yes, this is the key idea.
People use the same techniques to gain root access to a phone as they use to gain access to a system, or exploit a service. Think about low level stuff here (C/C++). The main idea is the same: confuse the faulty program and make it execute something that was not intended to be executed. Let's go over a short example in C, since your phone's kernel is written in that anyway.

Code:
#include <stdio.h> #include <string.h> void do_stuff(char **argv) { char buffer[80]; strcpy(buffer, argv[1]); printf("Your input was: %s\n", buffer); } int main(int argc, char **argv) { do_stuff(argv); return 0; }

Can you spot the bug here? On line 6, we copy argv[1] into buffer. buffer is 80 bytes long and argv[1] is as long as the user specifies. This is a clear example of a buffer overflow. What happens if we write too much data into buffer?
buffer is on the stack, so this will be a stack-based buffer overflow. The stack is also the place for key parts of the program's internals, like return addresses. Here is how a function is called in assembly:
Code:
main: ... call do_stuff ... do_stuff: ; do stuff here... ret

The call instruction will store the current instruction pointer (eip) on the stack, so when do_stuff is finished, it can issue a ret instruction. ret simply takes the first value from the stack and jumps to it.
So if we can write unrestricted amounts of data on the stack, we can overwrite this return address and jump to any parts of the program we want! Even better, we can jump to the data we just passed as a string, and execute it as code! The payload would look like this:

Code:
<assembly code> <padding to make the buffer reach the return address at the end> <address of our payload>

Things aren't this simple in real life though. They were in the 80s, but people caught on and implemented features that make the hacker's job a lot harder.
  • Sections have bits indicating whether they contain code or data. You can't execute code on a section that doesn't have the bit set (NX bit)
  • Sections have bits indicating whether they contain code or data. You can't execute code on a section that doesn't have the bit set (ASLR)
  • For stack-based stuff, they implemented canaries (also called cookies) - it's a random number on the stack that is checked before returning from each function - if the number changed, the program exits with a fault.


But hackers weren't lazy either, and they came up with awesome techniques to circumvent these restrictions.

To bypass the problem of randomized addresses and canaries, we usually need an info leak (think Heartbleed), which is basically one more bug to find in the code.

To bypass the NX bit, there is the Return-to-libc attack , which uses the code of the standard C library to gain access to a system. Since libc already has a system function that executes any OS commands, why do we need to write our own shellcode? We can just jump to the address of system with the parameters on the stack. It's not too hard to defend against this though - you can remove the sensitive functions from libc.


Here is where it starts to be mind-blowing to me. There is a technique called Return oriented programming (ROP), which is almost like the Return-to-libc attack, in the sense that we use the already existing code to gain access, but there's a catch! With ROP, we use tiny fractions of the code (a few bytes of bytecode at a time) and stitch the exploit together like that. There are lots of ret instructions in the available codebase, so we can just scan all of these and see if we have anything useful before the ret. Think about very basic stuff here, like moving a register's value to another register or taking a value from the stack. Since we control the values on the stack, we can put a lot of these ROP "gadgets" on the stack, and they will form a chain. The first one is executed, and then it returns to the next one, since its address is the next on the stack.

Quote:To bring it back to more to programming, how do the previous questions apply to hackers that crack software. Do they have to obtain the entire source code for a piece of software to somehow modify the code and recompile it? If so, how did they get the code in the first place?

This reply will address low level things again (C/C++). In a sense, they never obtain the source code. It depends on how you look at it, because in another sense the "source" is already there - it's the assembly code that was compiled from the high level language (I'm referring to high level here, because we are comparing it to assembly). All we need to do is reverse-engineer.

There are tools to disassemble the binary of the program we want to crack. IDA (Interactive Disassembler) is probably the most famous of these tools. The paid version even contains a decompiler. A decompiler is very different from a disassembler - it turns, or at least tries to turn, the assembly code back to C. Sometimes with surprisingly good results. The disassembler simply displays the bytecode in assembly form.

To crack a program, the assembly code is more than enough, the cracker doesn't need the source at all. You asked if the code was modified and recompiled. Almost - it's modified, but it isn't recompiled, since the source is not available. The modification is called patching. It consists of simply replacing byte code values in the binary.

As an example, probably the most simple protection mechanism would be to check for a serial number and jump to a routine that quits from the program if it's incorrect. Simply changing this jump instruction or the comparison instruction before is enough to "crack" a simple program like that.


All credits goes to "balidani" from his comment in this thread



[Image: V2NMLjp.png]


Reply

RE: How hackers exploit flaws in code ? #2
Added tutorial tag. Its more of a tutorial than resource.
[Image: OilyCostlyEwe.gif]

Reply

RE: How hackers exploit flaws in code ? #3
Very good thread man! I have enjoyed reading :Thumbs-Up:

Reply

RE: How hackers exploit flaws in code ? #4
Nice! I miss threads like this :Grin:
Fuck You.

Reply

RE: How hackers exploit flaws in code ? #5
Quote:Sections have bits indicating whether they contain code or data. You can't execute code on a section that doesn't have the bit set (ASLR)
I think this was a copy and paste error from the author; as that description is of NX/DEP not ASLR.
ASLR randomizes the address space making it more difficult to know the code address to jump to and such. It does so by randomizing the offset to various chunks of memory...the layout not specific parts of memory. Strings are still stored with other strings, code with code and stuff like that. So finding an information leak lets you calculate the offset to the memory you want.

Quote:The paid version even contains a decompiler.

This isn't true. IDA Pro just disassembles more architectures.

Hex-Rays Decompiler is a semi-seperate product from IDA. The Decompiler license is tied to an IDA license so its not totally separate.

Otherwise fairly accurate; high-level but accurate.

Reply