Showing posts with label Datatypes. Show all posts
Showing posts with label Datatypes. Show all posts

Format specifiers and input-output data management in C language

Format specifiers in C play a major role in storing or retrieving the data. Data may be corrupted or unexpected results may be produced if proper format specifiers are not used. Have a look at the list of format specifiers supported by C on gcc compiler.
Format specifiers in C on GCC
Format specifiers in C on GCC
Before going deep into format specifiers, two points are to be known and remembered throughout the input/output operations in C.

  1. Any input given does not directly stored into the memory. It is first moved to a buffer called ‘stdin’ and then the data fetched from the stdin to store into the memory (specifically RAM, when the process is running).
  2. Any output displayed on the console is not directly written on to the console by the data from the memory/processor. The data is first put into an intermediate buffer called ‘stdout’ and then fetched onto the console.

Data IO in C language
Data IO in C language
*Do not confuse with the term buffer. A buffer is simply a two-way data pool, acting as intermediary for data storage. The buffers stdin and stdout are never empty; they always consist of some data or the other, which we call garbage, unless explicitly defined by the user.
Format specifiers for scanning data from user:
There are different functions in C language supporting Input/output data streaming. Functions like scanf, gets, getc and getchar are used to collect data from the user. Except the scanf family of functions, all other functions for data input are predefined for particular datatype. Only scanf depends upon format specifiers to “scan” any type of data from the user, as shown above.
As we said earlier, whenever we input data to the computer, it is first stored in stdin buffer and then stored into the memory (RAM), after which the processor fetches data for further processing. But the question is…what is the amount of data fetched from the buffer? The format specifiers come into picture to define the amount of data to be fetched from the buffer to store in the memory.
For example, if %c is used, 1 byte of data from the stdin buffer is fetched; if %d or %i is used, 4 bytes of data is fetched from stdin. Consider the example shown below.
 #include<stdio.h>  
 main()  
 {  
  int f;  
  printf("Enter f:");  
  scanf("%c",&f);  
  printf("%d\t%c\n",f,f);  
 }  
This produces a warning after compilation; however, the program runs! When we enter an input, say a, the %d format specifier fetches a garbage value and %c fetches the exact value a.
Thought the variable f is declared as an integer. But we have used %c format specifier. This instructs the compiler to fetch 1 byte of data from stdin, where 4 bytes of memory is allocated for f. The 1 byte of data fetched from stdin is stored in the variable f for 4 bytes’ size. The remaining 3 bytes of data consist of some undefined values, resulting in garbage when %d is used to fetch 4 bytes of data.
Now consider the following example program.
 #include<stdio.h>  
 main()  
 {  
  char f;  
  printf("Enter f:");  
  scanf("%c",&f);  
  printf("%d\t*%c*\n",f,f);  
 }  
The variable f is declared with char datatype and the input is being stored into buffer with the format specifier %c. There is no mismatch in the datatype and the format specifier. So, this produces no warning and is compiled smoothly. After executing this program, you see certain output. But, this time, it is not a garbage. The value that %d and %c has fetched can be cross verified with – man ascii.
The same pattern of fetching data hold good for all the integral data types. The scene slightly differs for real datatypes. Execute the following program.
 #include<stdio.h>  
 main()  
 {  
  int f;  
  printf("Enter f:");  
  scanf("%d",&f);  
  printf("%d\t%f\n",f,f);  
 }  
Format specifiers for printing data onto console:
In the above example, though the variable f is declared as an integer, the format specifier changes how the data fetched from stdin is stored into memory. With the format specifier %f, the 4 bytes of data is fetched according to the IEEE 754 standard (click here to check floating format of IEEE 754 standard). But, when the data is fetched from the memory and displayed on console, the compiler just fetched 4 bytes normal of data, instead of fetching 4 bytes of data stored in IEEE format.
Suppose that the input provided during runtime is 5. When storing the data, it is stored by converting into IEEE format. But when fetching, irrespective of the bits stored, equivalent decimal data is fetched, which results as 1084227584 after binary to decimal conversion. Now, execute the following program and observe the difference.
 #include<stdio.h>  
 main()  
 {  
  int f;  
  printf("Enter f:");  
  scanf("%d",&f);  
  printf("%f\t%d\n",f,f);  
 }  
*Observe that we are using different format specifiers to fetch data available in the same variable.
According to the explanation given above, the format specifier %f should produce the equivalent output, for the bits stored by converting given input from decimal to binary. But, the output is shown as -0.000000. This is not because of any error; but because of the bug in the printf function itself. The printf function fails to fetch the data or corrupts the data when the format specifiers from both integral datatype and real datatype are used together on same variable and this cannot be avoided unless proper format specifier are used to store or fetch data. The same argument holds good from all the real datatypes.
 #include<stdio.h>  
 main()  
 {  
  int f;  
  printf("Enter f:");  
  scanf("%d",&f);  
  printf("%d\n",f);  
 }  
The bottom line of this post is:
The amount of data stored into the memory or fetched from the memory depends upon the format specifier used. On the other hand, the function printf and scanf fail to organise data from the user or data to the user when cross format specifiers from integral and real datatypes are used for the same variable.
Share:

Variation in data over the limits of datatypes

Every basic datatype is allocated with predefined size. Based upon the memory allocated, the range of values stored in a variable of particular datatype varies. Operations on the variables are smooth when they are within the lower and upper limits of these datatypes. It is ambiguous to guess the output when operations are being performed across the lower and upper limits, and vice versa.
Execute the following program and checkout the result.
 #include<stdio.h>  
 main()  
 {  
  char ch=127;  
  ch++;  
  printf("%d\n",ch);//printf("%c\n",ch);  
 }  
 #include<stdio.h>  
 main()  
 {  
  unsigned int i=-1;  
  i--;  
  printf("%d\n",i);  
 }  
Note that though ‘char’ is the datatype capable of storing characters and special symbols, internally it is considered as integer based upon the ASCII values and we can assign integer values to ‘char’ variables as shown. You can check the ASCII values in the manual page – man ascii
We know that gcc allocates 1byte of memory for a ‘char’ variable i.e., 8 bits. With these 8 bits, the char datatype is capable of storing 0-255 range of values. We know that the MSB of a byte is used for sign of the data stored in it. Hence, the range 0-255 (for unsigned char) changes from -128 to +127 (for signed char). All the ASCII values are stored from 0-127. Any value beyond this range produces unexpected results.
First, we have ch=127. Generally, when 127 is incremented by 1, it should be 128. But, the upper limit overflows to the lower limit as shown below. The same way, when the lower limit -128 is decremented by 1, it shows up with the value 128.
Data overflow in char datatype
Data overflow in char datatype
In the case of unsigned char, the lower limit becomes 0 and the upper limit becomes 255. The upper limit overflow navigates to the lower limit and the lower limit overflow results in the value shift to upper limit.
In simple words to say, the upper and lower limit overflow is circular – lower limit overflow results in a shift to upper limit and the upper limit over flow results in a shift to lower limit.
Till now, we illustrated how data overflows over the limits for ‘char’ datatype, which is 1byte in size. The same can be applied to integer also, with a catch that ‘int’ is of 4 bytes size in gcc. With 4 bytes, it is 32bits. Hence, the range of signed int is -216 to 216-1 and that of signed is 0 to 232-1. When the vaule 216-1 is incremented by 1, it moves to -216 and when the value -216 is decremented by 1, it moves to 216, for signed int. For unsigned int, 0 when decremented by 1 becomes 232-1 and when 232-1 is incremented by 1 becomes 0. Checkout the outputs of the following programs so that you can get a clear idea how data overflows at the boundaries.
 #include<stdio.h>  
 main()  
 {  
  int i=-1;//compiler takes int as signed int by default  
  i--;  
  printf("%d\n",i);  
 }  
Also, note that the same pattern of overflow occurs for all the integral datatypes.
Data overflow in int datatype
Data overflow in int datatype
The overflow pattern differs for real datatypes. Before that one needs know how float data is stored in the memory. IEEE 754 standard defined floating point storage. Every real data has 3 parts and they differ from float to double as shown below.

Floating point data storage
Floating point data storage
The keywords ‘signed’ and ‘unsigned’ are called sign qualifiers. There is no concept of sign qualifiers in real data. The sign is automatically stored in the MSB, as shown below.

IEEE 754 format for real data
IEEE 754 format for real data
How to convert float value to binary value?
Now, let us see how the values are stored in the sign, exponent and mantissa for a given real value. For example, let us convert 23.4 floating value to binary.
(23)10 = (10111)2
0.4 can be converted to binary as shown below.
0.4 x 2 = 0.8 -> 0
0.8 x 2 = 1.6 -> 1
1.6 x 2 = 1.2 -> 1
0.2 x 2 = 0.4 -> 0
0. 4 x 2 = 0.8 (recurring)
(0.4)10 = 0.01100110……..(recurring)
Now, (23.4)10 = (10111.01100110….)2
                           = 1.011101100110…..E+4
In the resultant binary representation, as the result is +ve, the sign bit is stored with 0, the bits after decimal point are stored in mantissa and +4 is the exponent. But, it is stored as 127+4, which is 10000011 in binary. Now, all these three parts are stored into memory as shown below.

Example for IEEE 754 floating point storage
Example for IEEE 754: 23.4 floating point storage in binary format
For double datatype, the exponent is stored as 1023+4.
increment or decrement operators show their effect only on the integral part of the floating point. For example, see the following program and check the output.
 #include<stdio.h>  
 main()  
 {  
  float f=1.23;  
  f++;  
  printf("%f\n",f);  
 }  
Overflow of data does not occur in real datatypes, unless it is intended to occur. So, there is no need to bother about how data varies over the limit in real datatypes, as we have infinite number of floating points between two given numbers.
Also, check the output of the program mentioned below, so that you can get a clear idea about floating points. Note that %e is used to display a floating point in exponential format.
 #include<stdio.h>  
 main()  
 {  
  float v=25.25;  
  printf("%f\n",v);  
  printf("%e\n",v);  
  v=12345.6;  
  printf("%f\n",v);  
  printf("%e\n",v);  
 }  
Share:

Datatypes in C language

All the programs we code and applications we develop are to for manipulations data. The way we store/fetch/process the data depends upon the datatype we use. The term datatype can be understood with two constraints.
The term ‘data’ can be interpreted in two ways – variable data and constant data. The term ‘type’ can be interpreted in two ways – integral type and real type.
There are two classifications of datatypes – basic/primitive datatypes and use user-defined datatypes. The chart shown below can make the difference between the interpretation of the term ‘datatype’ and classification of datatypes.
Datatypes in C
Datatypes in C
Interpretation of datatypes in C:
All the integral datatypes deal with the data that does not consist of decimal point. Though ‘char’ datatype deals with characters, gcc compiler considers ‘char’ as integer datatype. All the real datatypes deal with the data consisting of decimal point.
Irrespective of the data stored, data corresponding to a datatype can be constant or variable; throughout the program.
Classification of datatypes in C:
Size of the datatype, memory alignment and type of data to be store are the constraints for classification of datatypes.
The constraints of datatypes are known to the compiler in the basic datatypes.
For example, char occupies 1 byte of memory and an integer occupies 4 bytes of memory and so on.
The constraints of datatypes are known to the compiler only after compilation and/or sometimes when the program is being executed in the user-defined datatypes.
Execute the following program and observe the output.
 #include<stdio.h>  
 int main()  
 {  
  char v=50;  
  printf("v=%d\t%c\n",v,v);  
  v=v*2;  
  printf("v=%d\t%c\n",v,v);  
  v=v*2;  
  printf("v=%d\t%c\n",v,v);  
  return 0;  
 }  
Observe that the variable ‘v’ consists of the value 50, which is once fetched as an integer using %d and as a character using %c. When they are fetched using different format specifiers, the data is displayed according to ASCII values for %c.
The conclusion from this observation is that any type of data can be stored in any datatype but, the way the data is fetched and processed gives different results.
To get the datatypes, their explanation, ranges and format specifiers, click here.
*Remember that all the sizes of datatypes used in this blog are with respect to gcc compiler on 32-bit machine.  Also, note that the datatype Boolean is also added in the C99 version of C.
Share: